Run vLLM LLM inference tests with CVS#
2026-09-24
10 min read time
The vLLM suites measure LLM serving throughput, latency, and accuracy on AMD Instinct GPUs. vllm_single runs on the first cluster host and ignores additional hosts. vllm_distributed uses every host in the cluster file, with one-host fallback when only a single host is present. Packaged distributed recipes and thresholds are calibrated for two hosts; retune them before treating other sizes as pass/fail.
For a mapping-style node_dict, “first cluster host” is the first JSON key
in insertion order. vllm_single rewrites head_node_dict.mgmt_ip to that
host, so put the intended single-node target first in node_dict.
This page walks through a first run. For the full schema, every metric, and the threshold grammar, see vLLM inference benchmark configuration file for Cluster Validation Suite (CVS).
Prerequisites#
On every cluster node:
Docker installed, with the SSH user able to run it (passwordless
sudo dockeror membership in thedockergroup).Host driver loaded, so
/dev/kfd,/dev/dri/*, and/dev/infiniband/*are present for passthrough.A vLLM image either already loaded or pullable from a reachable registry. It must contain
vllmon the path — the suite invokesvllm serveandvllm bench serveinside the container.Model weights staged on a shared filesystem that every node mounts at the same path. Remote model download is not implemented, so weights must be present before the run.
A Hugging Face token file, if the model needs one. A pre-staged model can run without it.
On the head node where you launch cvs run:
CVS installed (see Install Cluster Validation Suite (CVS) on ROCm).
SSH key-based access to every cluster node.
For a multinode run, additionally have on hand:
The head node address reachable from every worker.
The network interface name used for inter-node traffic, for example
ens51f1np1. Find it withip -br addron a node. CVS cannot derive this automatically.
Step 1: Copy a configuration file#
CVS ships single-node and distributed vLLM configurations. List them:
cvs config list inference/vllm
Copy the one that matches your topology:
# Single node
cvs config copy inference/vllm/mi3xx_vllm_llama33-70b_fp8_single.json \
--output /tmp/cvs/vllm_singlenode_config.json
# Multiple nodes
cvs config copy inference/vllm/mi3xx_vllm_llama33-70b_fp8_distributed.json \
--output /tmp/cvs/vllm_multinode_config.json
You also need a cluster file describing your nodes. Use the container template, since the vLLM suite always runs inside a container:
cvs config copy cluster_container.json --output /tmp/cvs/cluster.json
Step 2: Fill in the placeholders#
Every value marked <changeme> must be replaced before the run.
In the cluster file, set your SSH user, private key path, and node addresses. See Run Cluster Validation Suite (CVS) test suites with a per-host Docker container backend for a walkthrough.
In the configuration file, set:
container.image— your vLLM image. Cite the full tag; do not abbreviate it.paths.shared_fs— the shared filesystem root. The other paths derive from it by default.paths.models_dir— where the weights live. Make sure this path is also mounted into the container bycontainer.runtime.args.volumes.paths.hf_token_file— path to your token file.server_params.model— the model to serve.
For a multinode configuration, also set:
container.env.NCCL_IB_HCA— the comma-separated HCA names exposed on every node. The packaged MI3xx configurations userdma0throughrdma7.container.env.NCCL_SOCKET_IFNAME— the Linux netdev associated with the selected RNICs.container.env.GLOO_SOCKET_IFNAMEandTP_SOCKET_IFNAME— generally the frontend/control-plane interface.container.env.NCCL_IB_GID_INDEX— the common fabric index reported byshow_gidson every selected HCA and node. Ifshow_gidsis unavailable, inspectibv_devinfo -vand the HCA’sgid_attrssysfs entries.
Tip
Leave enforce_thresholds set to false for your first run on new
hardware. The run reports every finite value produced from the 56-metric
registry without gating. Packaged threshold templates intentionally retain
the previous 32-metric assertion subset; you may add any other registered
metric or remove entries before enabling calibrated thresholds.
Step 3: Run the suite#
cvs run vllm_single \
--cluster_file /tmp/cvs/cluster.json \
--config_file /tmp/cvs/vllm_singlenode_config.json \
--html /tmp/cvs/vllm.html --self-contained-html \
--log-file /tmp/cvs/cvs.log
Use vllm_distributed for one distributed service across the cluster:
cvs run vllm_distributed \
--cluster_file /tmp/cvs/cluster.json \
--config_file /tmp/cvs/vllm_multinode_config.json \
--html /tmp/cvs/vllm.html --self-contained-html \
--log-file /tmp/cvs/cvs.log
Note
--self-contained-html only takes effect together with --html. Always pass both, so the report is a single file you can attach or copy off the cluster.
Any flag CVS does not recognize is passed straight through to pytest, so options such as -vvv and --capture=tee-sys work as usual.
Step 4: Read the results#
Open the HTML report. Each lifecycle stage, benchmark cell, and verification phase is its own row:
Lifecycle rows — container launch, topology discovery, model fetch, the OpenAI-compatible smoke test, then teardown. These tell you how far the run got.
Inference rows — one per sweep cell, labelled
<combo>-conc<N>.Verification rows — one per cell. Expand the row to see every finite metric plus every configured threshold. Only configured thresholds create subtests when enforcement is enabled. A missing or invalid gated value fails for every datasource, including GPU and Prometheus.
Results table — the summary is also printed to the console. Metric names are bare (for example,
output_throughputandqueue_time_p95_ms).
Per-cell logs land under your configured log_dir:
<log_dir>/vllm/out-node<rank>/isl<isl>_osl<osl>_conc<conc>/
vllm_serve_server.log <- the server's own log
client.log <- the load generator
results <- raw benchmark JSON
Important
When a run fails, read vllm_serve_server.log on the node, not just the CVS client log. Some faults — GPU exceptions, weight-loading failures, out-of-memory kills — appear only in the server log.
A skipped test_setup_sshd row is expected. vLLM communicates over the host network and needs no inter-container sshd.
Going multinode#
The cluster file determines the host count: vllm_distributed forms one service from every listed host, or falls back to a single-node run when only one host is present. For distributed runs, configure server_params.pipeline_parallel_size to match the intended layout and set the HCA, socket-interface, and GID settings under container.env. CVS derives the rendezvous address from the cluster head. Packaged configs use two hosts and pipeline_parallel_size of 2.
Using the default backend (mp)#
If you set nothing else, the suite uses mp. It requires pipeline parallelism across the nodes:
{
"container": {
"env": {
"NCCL_IB_HCA": "rdma0,rdma1,rdma2,rdma3,rdma4,rdma5,rdma6,rdma7",
"NCCL_SOCKET_IFNAME": "eno0",
"GLOO_SOCKET_IFNAME": "eno0",
"TP_SOCKET_IFNAME": "eno0",
"NCCL_IB_GID_INDEX": "3"
}
},
"server_params": {
"tensor_parallel_size": 8,
"pipeline_parallel_size": 2
}
}
CVS launches vllm serve on every node with the correct rank, adding --headless to every rank above 0.
Using ray#
Ray is opt-in. Set distributed_executor_backend in server_params:
{
"container": {
"env": {
"NCCL_IB_HCA": "rdma0,rdma1,rdma2,rdma3,rdma4,rdma5,rdma6,rdma7",
"NCCL_SOCKET_IFNAME": "eno0",
"GLOO_SOCKET_IFNAME": "eno0",
"TP_SOCKET_IFNAME": "eno0",
"NCCL_IB_GID_INDEX": "3"
}
},
"server_params": {
"tensor_parallel_size": 8,
"pipeline_parallel_size": 1,
"distributed_executor_backend": "ray"
}
}
CVS then bootstraps a Ray cluster before serving: ray start --head on rank 0, ray start --address=... on each worker, and ray stop at teardown. Only the head node runs vllm serve, so worker nodes produce no server log.
Note
Ray is not required for multinode — mp is the default and works across nodes. What ray changes is that it removes the pipeline-parallelism requirement, so pipeline_parallel_size of 1 becomes valid. Use ray when you want pure tensor-parallel serving across nodes; use the default otherwise.
Only the exact lowercase string "ray" selects it. Other values are rejected by configuration validation.
Common pitfalls#
The run fails immediately with a validation error. Configuration files are validated before anything launches, and every block except container rejects unknown keys — so a misspelled key is a hard error rather than a silently ignored setting. Read the message: it names the offending key.
Distributed topology validation fails. You configured multiple hosts on the default mp backend without pipeline parallelism. Either raise pipeline_parallel_size, or switch to ray.
“vllm_distributed requires container.env.NCCL_SOCKET_IFNAME”. Set all three socket-interface variables under container.env. The interface cannot be derived reliably from HCA names.
“Container image not specified in config”. container.image is empty. Watch for this specific trap: if your configuration file has a container block that omits image, it overwrites the cluster file’s image with an empty string. Set image in whichever file defines the block.
Container launch crashes with “too many values to unpack”. You placed env under container.runtime.args. It belongs at the container top level.
A threshold fails with “actual must be a finite built-in int or float”. The gated datasource did not produce a valid value. Inspect the raw benchmark artifact, GPU telemetry, or server metrics for that cell. The suite owns benchmark percentile collection; workload configuration cannot override it.
A verification parent skips. Either the benchmark produced no parseable result for that cell, threshold enforcement is disabled, or the cell has no active metric gates. Finite values are still retained as record rows.
A Run Deck baseline is incompatible. vLLM reports identify the bare metric
contract as {"id":"vllm-bare","version":1}. Historical reports containing
client.*, gpu.*, or prom.* metric keys cannot be compared. Generate
a new baseline with the current suite.
The sweep is slower than expected. Cells that differ only in concurrency reuse the running server; changing ISL, OSL, TP, PP, or any server argument forces a restart and a weight reload. Ordering runs so concurrency varies fastest avoids needless reloads.
Using runtime.name: "enroot". Enroot is registered but not implemented, and the run fails at container launch. Use docker.
See also#
vLLM inference benchmark configuration file for Cluster Validation Suite (CVS) — full configuration schema, metrics, and thresholds
Cluster Validation Suite (CVS) cluster file: configuration and backend selection — cluster file schema
Run Cluster Validation Suite (CVS) test suites with a per-host Docker container backend — container backend in depth
Run Cluster Validation Suite (CVS) test suites on AMD Instinct GPU clusters — running other CVS suites