Run vLLM LLM inference tests with CVS#

2026-09-24

10 min read time

Applies to Linux

The vLLM suites measure LLM serving throughput, latency, and accuracy on AMD Instinct GPUs. vllm_single runs on the first cluster host and ignores additional hosts. vllm_distributed uses every host in the cluster file, with one-host fallback when only a single host is present. Packaged distributed recipes and thresholds are calibrated for two hosts; retune them before treating other sizes as pass/fail.

For a mapping-style node_dict, “first cluster host” is the first JSON key in insertion order. vllm_single rewrites head_node_dict.mgmt_ip to that host, so put the intended single-node target first in node_dict.

This page walks through a first run. For the full schema, every metric, and the threshold grammar, see vLLM inference benchmark configuration file for Cluster Validation Suite (CVS).

Prerequisites#

On every cluster node:

  • Docker installed, with the SSH user able to run it (passwordless sudo docker or membership in the docker group).

  • Host driver loaded, so /dev/kfd, /dev/dri/*, and /dev/infiniband/* are present for passthrough.

  • A vLLM image either already loaded or pullable from a reachable registry. It must contain vllm on the path — the suite invokes vllm serve and vllm bench serve inside the container.

  • Model weights staged on a shared filesystem that every node mounts at the same path. Remote model download is not implemented, so weights must be present before the run.

  • A Hugging Face token file, if the model needs one. A pre-staged model can run without it.

On the head node where you launch cvs run:

For a multinode run, additionally have on hand:

  • The head node address reachable from every worker.

  • The network interface name used for inter-node traffic, for example ens51f1np1. Find it with ip -br addr on a node. CVS cannot derive this automatically.

Step 1: Copy a configuration file#

CVS ships single-node and distributed vLLM configurations. List them:

cvs config list inference/vllm

Copy the one that matches your topology:

# Single node
cvs config copy inference/vllm/mi3xx_vllm_llama33-70b_fp8_single.json \
  --output /tmp/cvs/vllm_singlenode_config.json

# Multiple nodes
cvs config copy inference/vllm/mi3xx_vllm_llama33-70b_fp8_distributed.json \
  --output /tmp/cvs/vllm_multinode_config.json

You also need a cluster file describing your nodes. Use the container template, since the vLLM suite always runs inside a container:

cvs config copy cluster_container.json --output /tmp/cvs/cluster.json

Step 2: Fill in the placeholders#

Every value marked <changeme> must be replaced before the run.

In the cluster file, set your SSH user, private key path, and node addresses. See Run Cluster Validation Suite (CVS) test suites with a per-host Docker container backend for a walkthrough.

In the configuration file, set:

  • container.image — your vLLM image. Cite the full tag; do not abbreviate it.

  • paths.shared_fs — the shared filesystem root. The other paths derive from it by default.

  • paths.models_dir — where the weights live. Make sure this path is also mounted into the container by container.runtime.args.volumes.

  • paths.hf_token_file — path to your token file.

  • server_params.model — the model to serve.

For a multinode configuration, also set:

  • container.env.NCCL_IB_HCA — the comma-separated HCA names exposed on every node. The packaged MI3xx configurations use rdma0 through rdma7.

  • container.env.NCCL_SOCKET_IFNAME — the Linux netdev associated with the selected RNICs.

  • container.env.GLOO_SOCKET_IFNAME and TP_SOCKET_IFNAME — generally the frontend/control-plane interface.

  • container.env.NCCL_IB_GID_INDEX — the common fabric index reported by show_gids on every selected HCA and node. If show_gids is unavailable, inspect ibv_devinfo -v and the HCA’s gid_attrs sysfs entries.

Tip

Leave enforce_thresholds set to false for your first run on new hardware. The run reports every finite value produced from the 56-metric registry without gating. Packaged threshold templates intentionally retain the previous 32-metric assertion subset; you may add any other registered metric or remove entries before enabling calibrated thresholds.

Step 3: Run the suite#

cvs run vllm_single \
  --cluster_file /tmp/cvs/cluster.json \
  --config_file /tmp/cvs/vllm_singlenode_config.json \
  --html /tmp/cvs/vllm.html --self-contained-html \
  --log-file /tmp/cvs/cvs.log

Use vllm_distributed for one distributed service across the cluster:

cvs run vllm_distributed \
  --cluster_file /tmp/cvs/cluster.json \
  --config_file /tmp/cvs/vllm_multinode_config.json \
  --html /tmp/cvs/vllm.html --self-contained-html \
  --log-file /tmp/cvs/cvs.log

Note

--self-contained-html only takes effect together with --html. Always pass both, so the report is a single file you can attach or copy off the cluster.

Any flag CVS does not recognize is passed straight through to pytest, so options such as -vvv and --capture=tee-sys work as usual.

Step 4: Read the results#

Open the HTML report. Each lifecycle stage, benchmark cell, and verification phase is its own row:

  • Lifecycle rows — container launch, topology discovery, model fetch, the OpenAI-compatible smoke test, then teardown. These tell you how far the run got.

  • Inference rows — one per sweep cell, labelled <combo>-conc<N>.

  • Verification rows — one per cell. Expand the row to see every finite metric plus every configured threshold. Only configured thresholds create subtests when enforcement is enabled. A missing or invalid gated value fails for every datasource, including GPU and Prometheus.

  • Results table — the summary is also printed to the console. Metric names are bare (for example, output_throughput and queue_time_p95_ms).

Per-cell logs land under your configured log_dir:

<log_dir>/vllm/out-node<rank>/isl<isl>_osl<osl>_conc<conc>/
  vllm_serve_server.log    <- the server's own log
  client.log               <- the load generator
  results                  <- raw benchmark JSON

Important

When a run fails, read vllm_serve_server.log on the node, not just the CVS client log. Some faults — GPU exceptions, weight-loading failures, out-of-memory kills — appear only in the server log.

A skipped test_setup_sshd row is expected. vLLM communicates over the host network and needs no inter-container sshd.

Going multinode#

The cluster file determines the host count: vllm_distributed forms one service from every listed host, or falls back to a single-node run when only one host is present. For distributed runs, configure server_params.pipeline_parallel_size to match the intended layout and set the HCA, socket-interface, and GID settings under container.env. CVS derives the rendezvous address from the cluster head. Packaged configs use two hosts and pipeline_parallel_size of 2.

Using the default backend (mp)#

If you set nothing else, the suite uses mp. It requires pipeline parallelism across the nodes:

{
  "container": {
    "env": {
      "NCCL_IB_HCA": "rdma0,rdma1,rdma2,rdma3,rdma4,rdma5,rdma6,rdma7",
      "NCCL_SOCKET_IFNAME": "eno0",
      "GLOO_SOCKET_IFNAME": "eno0",
      "TP_SOCKET_IFNAME": "eno0",
      "NCCL_IB_GID_INDEX": "3"
    }
  },
  "server_params": {
    "tensor_parallel_size": 8,
    "pipeline_parallel_size": 2
  }
}

CVS launches vllm serve on every node with the correct rank, adding --headless to every rank above 0.

Using ray#

Ray is opt-in. Set distributed_executor_backend in server_params:

{
  "container": {
    "env": {
      "NCCL_IB_HCA": "rdma0,rdma1,rdma2,rdma3,rdma4,rdma5,rdma6,rdma7",
      "NCCL_SOCKET_IFNAME": "eno0",
      "GLOO_SOCKET_IFNAME": "eno0",
      "TP_SOCKET_IFNAME": "eno0",
      "NCCL_IB_GID_INDEX": "3"
    }
  },
  "server_params": {
    "tensor_parallel_size": 8,
    "pipeline_parallel_size": 1,
    "distributed_executor_backend": "ray"
  }
}

CVS then bootstraps a Ray cluster before serving: ray start --head on rank 0, ray start --address=... on each worker, and ray stop at teardown. Only the head node runs vllm serve, so worker nodes produce no server log.

Note

Ray is not required for multinode — mp is the default and works across nodes. What ray changes is that it removes the pipeline-parallelism requirement, so pipeline_parallel_size of 1 becomes valid. Use ray when you want pure tensor-parallel serving across nodes; use the default otherwise.

Only the exact lowercase string "ray" selects it. Other values are rejected by configuration validation.

Common pitfalls#

The run fails immediately with a validation error. Configuration files are validated before anything launches, and every block except container rejects unknown keys — so a misspelled key is a hard error rather than a silently ignored setting. Read the message: it names the offending key.

Distributed topology validation fails. You configured multiple hosts on the default mp backend without pipeline parallelism. Either raise pipeline_parallel_size, or switch to ray.

“vllm_distributed requires container.env.NCCL_SOCKET_IFNAME”. Set all three socket-interface variables under container.env. The interface cannot be derived reliably from HCA names.

“Container image not specified in config”. container.image is empty. Watch for this specific trap: if your configuration file has a container block that omits image, it overwrites the cluster file’s image with an empty string. Set image in whichever file defines the block.

Container launch crashes with “too many values to unpack”. You placed env under container.runtime.args. It belongs at the container top level.

A threshold fails with “actual must be a finite built-in int or float”. The gated datasource did not produce a valid value. Inspect the raw benchmark artifact, GPU telemetry, or server metrics for that cell. The suite owns benchmark percentile collection; workload configuration cannot override it.

A verification parent skips. Either the benchmark produced no parseable result for that cell, threshold enforcement is disabled, or the cell has no active metric gates. Finite values are still retained as record rows.

A Run Deck baseline is incompatible. vLLM reports identify the bare metric contract as {"id":"vllm-bare","version":1}. Historical reports containing client.*, gpu.*, or prom.* metric keys cannot be compared. Generate a new baseline with the current suite.

The sweep is slower than expected. Cells that differ only in concurrency reuse the running server; changing ISL, OSL, TP, PP, or any server argument forces a restart and a weight reload. Ordering runs so concurrency varies fastest avoids needless reloads.

Using runtime.name: "enroot". Enroot is registered but not implemented, and the run fails at container launch. Use docker.

See also#