.. meta::
  :description: Run vLLM LLM inference tests with CVS on AMD Instinct GPUs, covering single-node and multinode distributed serving with ROCm and InfiniBand.
  :keywords: CVS, vLLM, inference, benchmark, AMD Instinct, ROCm, AMD, GPU, LLM, multinode, InfiniBand, RDMA

***************************************
Run vLLM LLM inference tests with CVS
***************************************

The vLLM suites measure LLM serving throughput, latency, and accuracy on AMD Instinct GPUs. ``vllm_single`` runs on the first cluster host and ignores additional hosts. ``vllm_distributed`` uses every host in the cluster file, with one-host fallback when only a single host is present. Packaged distributed recipes and thresholds are calibrated for two hosts; retune them before treating other sizes as pass/fail.

For a mapping-style ``node_dict``, "first cluster host" is the first JSON key
in insertion order. ``vllm_single`` rewrites ``head_node_dict.mgmt_ip`` to that
host, so put the intended single-node target first in ``node_dict``.

This page walks through a first run. For the full schema, every metric, and the threshold grammar, see :doc:`/reference/configuration-files/inference/vllm`.

Prerequisites
=============

On every cluster node:

- **Docker** installed, with the SSH user able to run it (passwordless ``sudo docker`` or membership in the ``docker`` group).
- **Host driver** loaded, so ``/dev/kfd``, ``/dev/dri/*``, and ``/dev/infiniband/*`` are present for passthrough.
- **A vLLM image** either already loaded or pullable from a reachable registry. It must contain ``vllm`` on the path — the suite invokes ``vllm serve`` and ``vllm bench serve`` inside the container.
- **Model weights** staged on a shared filesystem that every node mounts at the same path. Remote model download is not implemented, so weights must be present before the run.
- **A Hugging Face token file**, if the model needs one. A pre-staged model can run without it.

On the head node where you launch ``cvs run``:

- CVS installed (see :doc:`/install/install`).
- SSH key-based access to every cluster node.

For a multinode run, additionally have on hand:

- The **head node address** reachable from every worker.
- The **network interface name** used for inter-node traffic, for example ``ens51f1np1``. Find it with ``ip -br addr`` on a node. CVS cannot derive this automatically.

.. _vllm-set-up-config:

Step 1: Copy a configuration file
=================================

CVS ships single-node and distributed vLLM configurations. List them:

.. code:: bash

  cvs config list inference/vllm

Copy the one that matches your topology:

.. code:: bash

  # Single node
  cvs config copy inference/vllm/mi3xx_vllm_llama33-70b_fp8_single.json \
    --output /tmp/cvs/vllm_singlenode_config.json

  # Multiple nodes
  cvs config copy inference/vllm/mi3xx_vllm_llama33-70b_fp8_distributed.json \
    --output /tmp/cvs/vllm_multinode_config.json

You also need a cluster file describing your nodes. Use the container template, since the vLLM suite always runs inside a container:

.. code:: bash

  cvs config copy cluster_container.json --output /tmp/cvs/cluster.json

Step 2: Fill in the placeholders
================================

Every value marked ``<changeme>`` must be replaced before the run.

In the **cluster file**, set your SSH user, private key path, and node addresses. See :doc:`/how-to/run-with-containers` for a walkthrough.

In the **configuration file**, set:

- ``container.image`` — your vLLM image. Cite the full tag; do not abbreviate it.
- ``paths.shared_fs`` — the shared filesystem root. The other paths derive from it by default.
- ``paths.models_dir`` — where the weights live. Make sure this path is also mounted into the container by ``container.runtime.args.volumes``.
- ``paths.hf_token_file`` — path to your token file.
- ``server_params.model`` — the model to serve.

For a **multinode** configuration, also set:

- ``container.env.NCCL_IB_HCA`` — the comma-separated HCA names exposed on every node. The packaged MI3xx configurations use ``rdma0`` through ``rdma7``.
- ``container.env.NCCL_SOCKET_IFNAME`` — the Linux netdev associated with the selected RNICs.
- ``container.env.GLOO_SOCKET_IFNAME`` and ``TP_SOCKET_IFNAME`` — generally the frontend/control-plane interface.
- ``container.env.NCCL_IB_GID_INDEX`` — the common fabric index reported by ``show_gids`` on every selected HCA and node. If ``show_gids`` is unavailable, inspect ``ibv_devinfo -v`` and the HCA's ``gid_attrs`` sysfs entries.

.. tip::

  Leave ``enforce_thresholds`` set to ``false`` for your first run on new
  hardware. The run reports every finite value produced from the 56-metric
  registry without gating. Packaged threshold templates intentionally retain
  the previous 32-metric assertion subset; you may add any other registered
  metric or remove entries before enabling calibrated thresholds.

.. _vllm-run-tests:

Step 3: Run the suite
=====================

.. code:: bash

  cvs run vllm_single \
    --cluster_file /tmp/cvs/cluster.json \
    --config_file /tmp/cvs/vllm_singlenode_config.json \
    --html /tmp/cvs/vllm.html --self-contained-html \
    --log-file /tmp/cvs/cvs.log

Use ``vllm_distributed`` for one distributed service across the cluster:

.. code:: bash

  cvs run vllm_distributed \
    --cluster_file /tmp/cvs/cluster.json \
    --config_file /tmp/cvs/vllm_multinode_config.json \
    --html /tmp/cvs/vllm.html --self-contained-html \
    --log-file /tmp/cvs/cvs.log

.. note::

  ``--self-contained-html`` only takes effect together with ``--html``. Always pass both, so the report is a single file you can attach or copy off the cluster.

  Any flag CVS does not recognize is passed straight through to pytest, so options such as ``-vvv`` and ``--capture=tee-sys`` work as usual.

Step 4: Read the results
========================

Open the HTML report. Each lifecycle stage, benchmark cell, and verification phase is its own row:

- **Lifecycle rows** — container launch, topology discovery, model fetch, the OpenAI-compatible smoke test, then teardown. These tell you *how far* the run got.
- **Inference rows** — one per sweep cell, labelled ``<combo>-conc<N>``.
- **Verification rows** — one per cell. Expand the row to see every finite metric plus every configured threshold. Only configured thresholds create subtests when enforcement is enabled. A missing or invalid gated value fails for every datasource, including GPU and Prometheus.
- **Results table** — the summary is also printed to the console. Metric names are bare (for example, ``output_throughput`` and ``queue_time_p95_ms``).

Per-cell logs land under your configured ``log_dir``::

  <log_dir>/vllm/out-node<rank>/isl<isl>_osl<osl>_conc<conc>/
    vllm_serve_server.log    <- the server's own log
    client.log               <- the load generator
    results                  <- raw benchmark JSON

.. important::

  When a run fails, read ``vllm_serve_server.log`` on the node, not just the CVS client log. Some faults — GPU exceptions, weight-loading failures, out-of-memory kills — appear only in the server log.

A skipped ``test_setup_sshd`` row is expected. vLLM communicates over the host network and needs no inter-container sshd.

Going multinode
===============

The cluster file determines the host count: ``vllm_distributed`` forms one service from every listed host, or falls back to a single-node run when only one host is present. For distributed runs, configure ``server_params.pipeline_parallel_size`` to match the intended layout and set the HCA, socket-interface, and GID settings under ``container.env``. CVS derives the rendezvous address from the cluster head. Packaged configs use two hosts and ``pipeline_parallel_size`` of 2.

Using the default backend (mp)
------------------------------

If you set nothing else, the suite uses ``mp``. It requires pipeline parallelism across the nodes:

.. code:: json

    {
      "container": {
        "env": {
          "NCCL_IB_HCA": "rdma0,rdma1,rdma2,rdma3,rdma4,rdma5,rdma6,rdma7",
          "NCCL_SOCKET_IFNAME": "eno0",
          "GLOO_SOCKET_IFNAME": "eno0",
          "TP_SOCKET_IFNAME": "eno0",
          "NCCL_IB_GID_INDEX": "3"
        }
      },
      "server_params": {
        "tensor_parallel_size": 8,
        "pipeline_parallel_size": 2
      }
    }

CVS launches ``vllm serve`` on every node with the correct rank, adding ``--headless`` to every rank above 0.

Using ray
---------

Ray is opt-in. Set ``distributed_executor_backend`` in ``server_params``:

.. code:: json

    {
      "container": {
        "env": {
          "NCCL_IB_HCA": "rdma0,rdma1,rdma2,rdma3,rdma4,rdma5,rdma6,rdma7",
          "NCCL_SOCKET_IFNAME": "eno0",
          "GLOO_SOCKET_IFNAME": "eno0",
          "TP_SOCKET_IFNAME": "eno0",
          "NCCL_IB_GID_INDEX": "3"
        }
      },
      "server_params": {
        "tensor_parallel_size": 8,
        "pipeline_parallel_size": 1,
        "distributed_executor_backend": "ray"
      }
    }

CVS then bootstraps a Ray cluster before serving: ``ray start --head`` on rank 0, ``ray start --address=...`` on each worker, and ``ray stop`` at teardown. Only the head node runs ``vllm serve``, so worker nodes produce no server log.

.. note::

  Ray is not required for multinode — ``mp`` is the default and works across nodes. What ray changes is that it **removes the pipeline-parallelism requirement**, so ``pipeline_parallel_size`` of 1 becomes valid. Use ray when you want pure tensor-parallel serving across nodes; use the default otherwise.

  Only the exact lowercase string ``"ray"`` selects it. Other values are rejected by configuration validation.

Common pitfalls
===============

**The run fails immediately with a validation error.** Configuration files are validated before anything launches, and every block except ``container`` rejects unknown keys — so a misspelled key is a hard error rather than a silently ignored setting. Read the message: it names the offending key.

**Distributed topology validation fails.** You configured multiple hosts on the default ``mp`` backend without pipeline parallelism. Either raise ``pipeline_parallel_size``, or switch to ray.

**"vllm_distributed requires container.env.NCCL_SOCKET_IFNAME".** Set all three socket-interface variables under ``container.env``. The interface cannot be derived reliably from HCA names.

**"Container image not specified in config".** ``container.image`` is empty. Watch for this specific trap: if your configuration file has a ``container`` block that omits ``image``, it overwrites the cluster file's image with an empty string. Set ``image`` in whichever file defines the block.

**Container launch crashes with "too many values to unpack".** You placed ``env`` under ``container.runtime.args``. It belongs at the ``container`` top level.

**A threshold fails with "actual must be a finite built-in int or float".** The gated datasource did not produce a valid value. Inspect the raw benchmark artifact, GPU telemetry, or server metrics for that cell. The suite owns benchmark percentile collection; workload configuration cannot override it.

**A verification parent skips.** Either the benchmark produced no parseable result for that cell, threshold enforcement is disabled, or the cell has no active metric gates. Finite values are still retained as record rows.

**A Run Deck baseline is incompatible.** vLLM reports identify the bare metric
contract as ``{"id":"vllm-bare","version":1}``. Historical reports containing
``client.*``, ``gpu.*``, or ``prom.*`` metric keys cannot be compared. Generate
a new baseline with the current suite.

**The sweep is slower than expected.** Cells that differ only in concurrency reuse the running server; changing ISL, OSL, TP, PP, or any server argument forces a restart and a weight reload. Ordering ``runs`` so concurrency varies fastest avoids needless reloads.

**Using** ``runtime.name: "enroot"``. Enroot is registered but not implemented, and the run fails at container launch. Use ``docker``.

See also
========

- :doc:`/reference/configuration-files/inference/vllm` — full configuration schema, metrics, and thresholds
- :doc:`/reference/cluster/cluster-file` — cluster file schema
- :doc:`/how-to/run-with-containers` — container backend in depth
- :doc:`/how-to/test-suites/index` — running other CVS suites
