.. meta::
  :description: Reference for the vLLM inference benchmark configuration in CVS, covering container setup, sweep cells, threshold files, accuracy tests, and multi-node execution.
  :keywords: CVS, vLLM, inference, ROCm, LLM, benchmark, GPU, AMD, threshold, multinode, accuracy, JSON, configuration

****************************************************************************
vLLM inference benchmark configuration file for Cluster Validation Suite (CVS)
****************************************************************************

The vLLM suites benchmark LLM serving throughput, latency, and accuracy on AMD Instinct GPUs. ``vllm_single`` runs on the first cluster host and ignores additional hosts. ``vllm_distributed`` uses every host in the cluster file, with one-host fallback when only a single host is present. Packaged distributed recipes and thresholds are calibrated for two hosts; retune them before treating other sizes as pass/fail.

For a mapping-style ``node_dict``, "first cluster host" means the first JSON key
in insertion order. ``vllm_single`` scopes execution to that host and rewrites
``head_node_dict.mgmt_ip`` to match, even if the original head names another
node. Put the intended single-node target first in ``node_dict``.

Run it with:

.. code:: bash

  cvs run vllm_single --cluster_file <cluster.json> --config_file <config.json>

For a step-by-step walkthrough of a first run, see :doc:`/how-to/test-suites/inference/vllm`. This page is the schema and metric reference.

Lifecycle
=========

Each stage of the run is an independent test, so every stage becomes its own timed, pass/fail row in the HTML report. The suite pins this order explicitly rather than relying on definition order:

.. list-table::
   :widths: 1 3 6
   :header-rows: 1

   * - Order
     - Test
     - Purpose
   * - 0
     - ``test_launch_container``
     - Pull/load the image and start the container on every node.
   * - 1
     - ``test_setup_sshd``
     - Always skipped for vLLM (see note below).
   * - 2
     - ``test_discover_topology``
     - Resolve IB HCA devices; no-op for effective single-node execution.
   * - 3
     - ``test_model_fetch``
     - Stage model weights.
   * - 4
     - ``test_openai_compatible_smoke``
     - Short-lived server; verifies the OpenAI-compatible API answers.
   * - 5
     - ``test_vllm_inference``
     - Run one benchmark cell (parametrized per sweep run).
   * - 6
     - ``test_verify_cell_metrics``
     - One verification parent per cell; configured threshold gates are listed as subtests.
   * - 7
     - ``test_accuracy_eval``
     - lm-eval accuracy tasks, if any are configured.
   * - 8
     - ``test_print_results_table``
     - Console + report summary table.
   * - 9
     - ``test_teardown``
     - Stop the server and tear down the container.

.. note::

  ``test_setup_sshd`` always skips in this suite. vLLM uses ``--distributed-executor-backend mp`` with NCCL over the host network, so no inter-container sshd is needed. A skipped row here is expected, not a problem.

.. note::

  vLLM suite execution is serial and single-pass. Do not use xdist workers or
  ``pytest-repeat`` counts above one: the verification phase consumes results
  collected earlier in the same pytest process.

If a stage fails, later stages are skipped rather than cascading into confusing downstream errors. The container is still torn down by a leak-guard even when a mid-sweep test fails.

Configuration file structure
============================

A vLLM configuration file has these top-level keys:

.. list-table::
   :widths: 3 2 5
   :header-rows: 1

   * - Key
     - Required
     - Description
   * - ``enforce_thresholds``
     - no (default ``true``)
     - When ``false``, metrics record without requiring calibrated threshold cells.
   * - ``threshold_json``
     - yes
     - Explicit path to the threshold file. See :ref:`vllm-threshold-discovery`.
   * - ``container``
     - yes
     - Container/Docker settings. See :ref:`vllm-container`.
   * - ``paths``
     - yes
     - Filesystem locations. See :ref:`vllm-paths`.
   * - ``server_params``
     - yes
     - Harness-owned server fields plus snake-case ``vllm serve`` options.
   * - ``benchmark_params``
     - no
     - Benchmark defaults plus snake-case ``vllm bench serve`` options.
   * - ``sweeps``
     - yes
     - Canonical run-cell keys mapped to benchmark overrides.
   * - ``runs``
     - yes
     - Nonempty ordered list of the sweep cells to execute.
   * - ``thresholds``
     - no
     - Per-cell pass/fail specs. See :ref:`vllm-thresholds`.
   * - ``accuracy``
     - no
     - lm-eval task selection. See :ref:`vllm-accuracy`.

.. important::

  Top-level and structural blocks **forbid unknown keys**. ``server_params``,
  ``benchmark_params``, and individual sweep overrides deliberately accept
  arbitrary snake-case vLLM option names. CVS converts them to kebab-case CLI
  flags. ``null`` omits a flag, ``true`` emits a bare flag, scalars emit one
  value, lists emit one flag followed by values, and mappings emit compact JSON.
  Use an option's negative form instead of ``false``.

Placeholder substitution
------------------------

Values are resolved in three passes, so later forms can reference earlier ones:

1. **Cluster placeholders** — ``{user-id}`` resolves to the current username.
2. **Self-reference within** ``paths`` — ``{shared_fs}`` expands to the already-resolved ``paths.shared_fs``.
3. **Cross-block** — ``{paths.models_dir}`` expands anywhere else in the file, such as in a volume mount.

.. code:: json

    {
      "paths": {
        "shared_fs": "/mnt/dtni/{user-id}",
        "models_dir": "{shared_fs}/models",
        "log_dir": "{shared_fs}/LOGS",
        "hf_token_file": "{shared_fs}/.cache/huggingface/token"
      },
      "container": {
        "runtime": {
          "args": {
            "volumes": ["{paths.models_dir}:/models"]
          }
        }
      }
    }

.. _vllm-backends:

Execution backends
==================

Four different things in this stack are called a "backend". They are unrelated, and confusing them is the most common configuration mistake.

.. list-table::
   :widths: 3 2 5
   :header-rows: 1

   * - Setting
     - Values
     - What it selects
   * - ``benchmark_params.backend``
     - ``"vllm"`` (default)
     - The **client** backend passed to ``vllm bench serve --backend``. Nothing to do with distribution.
   * - ``server_params.distributed_executor_backend``
     - ``"mp"`` (default), ``"ray"``
     - How vLLM distributes the model across nodes. This is the multinode setting.
   * - ``container.runtime.name``
     - ``"docker"``
     - The container runtime.
   * - Cluster file ``orchestrator``
     - ``"baremetal"``, ``"container"``
     - Whether CVS runs commands on the host or inside a container. See :doc:`/reference/cluster/cluster-file`.

.. warning::

  ``enroot`` is registered but **not implemented**. Every method is a stub that returns failure, so a run with ``runtime.name: "enroot"`` fails at ``test_launch_container``. Podman is not supported. Use ``docker``.

Distributed executor: mp and ray
--------------------------------

Multinode runs support **two** executor backends. ``mp`` is the default and requires no configuration key at all.

**mp (default).** Used whenever ``server_params.distributed_executor_backend`` is absent. The suite injects the full distributed block into each rank's ``vllm serve`` command:

.. code:: bash

  vllm serve <model> --tensor-parallel-size <tp> --port <port> \
    --node-rank <rank> --master-addr <addr> --master-port <port> \
    --nnodes <n> --pipeline-parallel-size <pp> \
    --distributed-executor-backend mp

Every rank above 0 additionally gets ``--headless``. This path **requires pipeline parallelism** (``pipeline_parallel_size`` greater than 1).

**ray (opt-in).** Selected by setting ``server_params.distributed_executor_backend`` to the exact lowercase string ``"ray"``. Other values are configuration errors. Ray takes a completely different route:

1. Bootstrap the cluster head: ``ray start --head --port=<master_port>``
2. Bootstrap each worker: ``ray start --address=<cluster-head>:<dist_init_port>``
3. Launch ``vllm serve`` on the **head node only** — workers run no serve process
4. On teardown, broadcast ``ray stop`` after the process kill

Under ray, none of the mp distributed flags are emitted. ``--pipeline-parallel-size`` is added only when ``server_params.pipeline_parallel_size`` is greater than 1.

.. note::

  Ray does not *enable* multinode — it **relaxes** the pipeline-parallelism requirement. With ray, ``pipeline_parallel_size`` of 1 is legal and is the expected configuration for pure tensor-parallel multinode serving. With mp, pipeline parallelism is mandatory.

Because only the head node serves under ray, worker ranks produce no per-rank server log. That is expected.

Topology validation rules
-------------------------

These rules are enforced when the configuration file loads, before anything starts:

.. list-table::
   :widths: 4 6
   :header-rows: 1

   * - Condition
     - Rule
   * - two or more cluster hosts, backend is not ray
     - ``pipeline_parallel_size`` **must** be greater than 1
   * - two or more cluster hosts, backend is ray
     - ``pipeline_parallel_size`` of 1 is valid
   * - ``pipeline_parallel_size`` > 1
     - The distributed suite requires more than one cluster host
   * - two or more cluster hosts, either backend
     - ``container.env.NCCL_SOCKET_IFNAME`` is **required**

The corresponding error messages are:

.. code:: text

  multi-host distributed execution requires pipeline_parallel_size > 1 unless using ray
  pipeline_parallel_size > 1 requires a multi-host distributed suite
  vllm_distributed requires container.env.NCCL_SOCKET_IFNAME on multi-host clusters

Multinode prerequisites
-----------------------

Beyond the validation rules, a multinode run needs:

- ``server_params.dist_init_port`` — default ``29501``; CVS derives the head address from the cluster.
- ``container.env.NCCL_IB_HCA`` — the comma-separated RDMA HCA names available on every node. The packaged MI3xx configurations set ``rdma0`` through ``rdma7``.
- ``container.env.NCCL_SOCKET_IFNAME`` — the Linux netdev associated with the selected RNICs.
- ``container.env.GLOO_SOCKET_IFNAME`` and ``TP_SOCKET_IFNAME`` — generally the frontend/control-plane interface.
- ``container.env.NCCL_IB_GID_INDEX`` — the index for the intended RoCE/IB fabric. Use ``show_gids`` inside the container and choose an entry available on every selected HCA and node. If that command is unavailable, inspect ``ibv_devinfo -v`` and ``/sys/class/infiniband/<hca>/ports/<port>/gid_attrs/``.

.. _vllm-container:

Container and Docker configuration
==================================

The container block controls image selection, lifetime, and the ``docker run`` flags.

.. code:: json

    {
      "container": {
        "lifetime": "per_run",
        "name": "vllm_perf_inference_rocm",
        "image": "rocm/vllm:latest",
        "runtime": {
          "name": "docker",
          "args": {
            "network": "host",
            "ipc": "host",
            "privileged": true,
            "volumes": [
              "/home/{user-id}:/home/{user-id}",
              "{paths.models_dir}:/models"
            ]
          }
        }
      }
    }

Container block keys
--------------------

The following keys are accepted inside the ``container`` block.

.. list-table::
   :widths: 3 2 5
   :header-rows: 1

   * - Key
     - Default
     - Description
   * - ``image``
     - none
     - Container image. **Required** — launch fails with ``Container image not specified in config``.
   * - ``name``
     - ``<user>_<sanitized-image>``
     - Container name.
   * - ``lifetime``
     - ``"per_run"``
     - One of ``no_launch``, ``per_run``, ``persistent``.
   * - ``runtime.name``
     - ``"docker"``
     - Container runtime.
   * - ``runtime.args``
     - ``{}``
     - Docker flags; see the table below.
   * - ``env``
     - ``{}``
     - Container-level environment variables. **Top level, not under** ``runtime.args``.
   * - ``image_tar``
     - absent
     - Path on each host to a saved image tar to ``docker load`` instead of pulling. **Top level**.

.. warning::

  Put ``env`` at the **container top level**. Placing it under ``runtime.args`` crashes container launch: the code iterates that value as a sequence of pairs, which raises ``ValueError: too many values to unpack`` for any key longer than two characters.

Runtime arguments
-----------------

All keys under ``runtime.args`` are optional. **List-valued keys append to the defaults; scalar keys override them.** There is no way to remove a default device or capability.

.. list-table::
   :widths: 2 1 3 4
   :header-rows: 1

   * - Key
     - Merge
     - Default
     - Emitted flag
   * - ``volumes``
     - append
     - ``/home/$USER/.ssh:/host_ssh`` (always added)
     - ``-v <host>:<ctr>[:ro]``
   * - ``devices``
     - append
     - ``/dev/kfd``, ``/dev/dri``, ``/dev/infiniband``
     - ``--device <path>``
   * - ``cap_add``
     - append
     - ``SYS_PTRACE``, ``IPC_LOCK``, ``SYS_ADMIN``
     - ``--cap-add <cap>``
   * - ``security_opt``
     - append
     - ``seccomp=unconfined``, ``apparmor=unconfined``
     - ``--security-opt <opt>``
   * - ``group_add``
     - append
     - ``video``
     - ``--group-add <group>``
   * - ``ulimit``
     - append
     - ``memlock=-1``
     - ``--ulimit <limit>``
   * - ``network``
     - override
     - ``host``
     - ``--network <mode>``
   * - ``ipc``
     - override
     - ``host``
     - ``--ipc <mode>``
   * - ``privileged``
     - override
     - ``true``
     - ``--privileged``
   * - ``registry``
     - n/a
     - none
     - Triggers ``docker login``; see below

The assembled command is:

.. code:: bash

  docker run -d --name <name> <args> <image> sleep infinity

The container is a long-lived sidecar; every workload command runs through ``docker exec`` inside it. InfiniBand devices are additionally passed through by per-host shell expansion at launch time, so each node mounts the devices it actually has.

.. note::

  ``--gpus`` is deliberately never emitted — GPU access on AMD hardware comes from the ``/dev/kfd`` and ``/dev/dri`` device mounts plus the ``video`` group.

  ``shm_size`` is **not supported** on this path. Setting ``runtime.args.shm_size`` is silently ignored; ``--shm-size`` is never emitted.

Container lifetime
------------------

The ``lifetime`` key controls when the container is started and stopped relative to the test lifecycle.

.. list-table::
   :widths: 2 4 4
   :header-rows: 1

   * - ``lifetime``
     - Setup behavior
     - Teardown behavior
   * - ``no_launch``
     - Verifies a container of that name is already running; never starts one.
     - No-op
   * - ``per_run``
     - Force-removes any stale container of the same name, then launches.
     - ``docker rm -f``
   * - ``persistent``
     - Attaches if running on all hosts; cold-starts if absent on all hosts; **refuses** on partial or failed probe.
     - No-op

.. tip::

  With ``persistent``, always pin ``container.name`` explicitly. The default name is derived from the image, so bumping an image tag silently abandons the old container and starts a new one.

Registry authentication
-----------------------

Set ``runtime.args.registry`` to log in before pulling:

.. code:: json

    {
      "registry": {
        "username": "myuser",
        "password_file": "/path/on/each/host/to/token",
        "server": "registry.example.com"
      }
    }

``username`` and ``password_file`` are required; ``server`` defaults to Docker Hub. The password is read from a **file path on each remote host** — there is no inline password or token key, and the login is kept out of the logs. Login is skipped entirely when ``image_tar`` is set, since a tar load never pulls.

Image resolution order at launch:

1. If ``image_tar`` is set and the image is absent, ``docker load`` it.
2. Otherwise, if ``registry`` is set, log in.
3. Check whether the image exists on **all** hosts.
4. If not, ``docker pull`` it, with no retry or fallback.

.. note::

  The image-exists check matches ``Repository:Tag`` exactly, so an image referenced without a tag or by digest never matches and is pulled on every run.

Cluster file merge
------------------

The variant's ``container`` block is deep-merged **onto** the cluster file's block: dictionaries merge key-wise, while scalars and lists are replaced. Cluster-set values survive unless the variant sets the same key.

.. warning::

  A ``container`` block in the variant always contributes ``lifetime``, ``name``, and ``image`` — including empty defaults. If your variant defines ``container`` but omits ``image``, it overwrites a cluster-file ``image`` with an empty string and the launch fails. Set ``image`` in whichever file defines the block.

.. _vllm-paths:

Paths
=====

All four keys are required.

.. list-table::
   :widths: 3 7
   :header-rows: 1

   * - Key
     - Description
   * - ``shared_fs``
     - Root of the shared filesystem, typically the anchor other paths reference.
   * - ``models_dir``
     - Model weight cache; exported into the server as ``HF_HUB_CACHE``.
   * - ``log_dir``
     - Root for run artifacts.
   * - ``hf_token_file``
     - Path to a file containing the Hugging Face token.

If ``hf_token_file`` does not exist, the run continues with an empty token. vLLM
configs always serve a pre-staged model mounted under ``paths.models_dir``.

Per-cell artifacts land in::

  <log_dir>/vllm/out-node<rank>/isl<isl>_osl<osl>_conc<conc>/
    vllm_serve_server.log
    client.log
    results

.. _vllm-model:

Model
=====

.. list-table::
   :widths: 3 2 5
   :header-rows: 1

   * - Key
     - Default
     - Description
   * - ``server_params.model``
     - none
     - Local path, for example ``/models/Llama-3.1-70B-Instruct-FP8-KV``

.. _vllm-roles:

Server role
===========

``server_params`` controls the ``vllm serve`` process. Harness-owned fields
are ``model``, ``tensor_parallel_size``, ``pipeline_parallel_size``, ``port``,
``dist_init_port``, polling controls, and ``distributed_executor_backend``.
Every other snake-case key is passed through to ``vllm serve``.

.. list-table::
   :widths: 3 2 5
   :header-rows: 1

   * - Key
     - Default
     - Description
   * - ``model``
     - none
     - Local model path supplied as the positional ``vllm serve`` argument.
   * - ``tensor_parallel_size``
     - none
     - Tensor-parallel degree.
   * - ``pipeline_parallel_size``
     - ``1``
     - Pipeline-parallel degree.
   * - ``port``
     - ``8888``
     - OpenAI-compatible server port.

How server_params are flattened
-------------------------------

.. list-table::
   :widths: 3 3 4
   :header-rows: 1

   * - JSON value
     - Emitted
     - Example
   * - Scalar
     - ``--flag value``
     - ``"kv_cache_dtype": "fp8"`` → ``--kv-cache-dtype fp8``
   * - ``true``
     - ``--flag`` (bare)
     - ``"enforce_eager": true`` → ``--enforce-eager``
   * - ``false``
     - rejected
     - Omit the setting or use vLLM's explicit negative option.
   * - List
     - One flag followed by its values.
     - ``"x": ["a","b"]`` → ``--x a b``

When ``server_params.max_model_len`` is absent, CVS derives one for each
benchmark cell from its effective parameters after sweep overrides:

.. code:: text

  ceil((ISL + OSL) * (1 + random_range_ratio)) + random_prefix_len + 8

An explicit non-null ``server_params.max_model_len`` takes precedence and
emits exactly one ``--max-model-len`` flag. An explicit null value is an
intentional opt-out: the generic option serializer emits no flag, and CVS
suppresses the fallback so vLLM uses the model or image default.

Environment variables: two mechanisms
-------------------------------------

These are separate and are frequently confused.

.. list-table::
   :widths: 2 4 4
   :header-rows: 1

   * -
     - ``container.env``
     - generated per-command environment
   * - Applied by
     - ``docker run -e``
     - A sourced shell script inside the container after HCA discovery.
   * - Scope
     - Every command in the container, for its whole lifetime.
     - Hugging Face path/token variables and legacy network fallbacks.
   * - Changing it
     - Requires recreating the container.
     - Takes effect on the next command.
   * - Defaults
     - ``GPUS=8``, ``MULTINODE=true``
     - No network overrides unless a legacy top-level field is set.

The generated per-command environment always exports:

.. code:: bash

  export HF_TOKEN=<token>
  export HF_HUB_CACHE=<paths.models_dir>

Packaged configurations set the HCA, socket-interface, GID, and NCCL debug
settings in ``container.env`` so every command inherits them and topology
discovery does not overwrite them. The legacy top-level ``ib_hca_devices`` and
``ib_netdev`` fields remain fallbacks for external configurations that omit
their corresponding container environment variables. Put other static ROCm,
NCCL, and vLLM exports in ``container.env``.

The packaged MI3xx catalog's AITER settings are image- and model-scoped. For
the image based on vLLM commit ``4bdc8a788``:

- DeepSeek V4 Flash, DeepSeek V4 Pro, GLM 5.1, and GLM 5.2 set
  ``VLLM_ROCM_USE_AITER=1``, ``VLLM_ROCM_USE_AITER_MHA=0``, and
  ``GPU_ARCHS=gfx942``.
- Kimi K2.5 sets ``VLLM_ROCM_USE_AITER=1`` and disables
  ``VLLM_ROCM_USE_AITER_MHA``, ``VLLM_ROCM_USE_AITER_FP4BMM``,
  ``VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS``, and
  ``VLLM_ROCM_USE_AITER_MLA``.
- Other packaged model families carry no AITER overrides.

Do not add ``VLLM_USE_AITER_UNIFIED_ATTENTION`` or
``VLLM_ROCM_USE_AITER_FUSED_MOE_A16W4``: this image does not register them, so
they are warning-only and ignored. Do not substitute
``VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION`` or other attention-backend variables
without validating both registration and call sites in the exact image.

.. _vllm-params:

Benchmark parameters
====================

``benchmark_params`` holds client defaults. Per-cell entries in ``sweeps``
override these values. CVS owns endpoint construction, result paths, and
percentile reporting.

.. list-table::
   :widths: 3 2 5
   :header-rows: 1

   * - Key
     - Default
     - Description
   * - ``backend``
     - ``"vllm"``
     - Client backend for ``vllm bench serve``.
   * - ``base_url``
     - ``"http://0.0.0.0"``
     - Server base URL.
   * - ``num_prompts``
     - ``3200``
     - Total prompts per cell.
   * - ``dataset_name``
     - ``"random"``
     - Dataset for the load generator.
   * - ``num_prompts``
     - ``"3200"``
     - Total prompts per cell.
   * - ``burstiness``
     - ``"1.0"``
     - 1.0 is a uniform arrival process; lower is burstier.
   * - ``seed``
     - ``"0"``
     - Random seed.
   * - ``request_rate``
     - ``"inf"``
     - Arrival rate; ``inf`` sends as fast as concurrency allows.
   * - ``random_range_ratio``
     - ``"0.0"``
     - Length jitter around ISL/OSL; also feeds the derived max-model-len.
   * - ``random_prefix_len``
     - ``"0"``
     - Shared prefix length.
   * - ``tokenizer_mode``
     - ``"auto"``
     - Tokenizer mode.
   * - ``client_poll_iterations``
     - ``20``
     - Client completion polls before giving up.

.. tip::

  Arbitrary snake-case keys under ``benchmark_params`` and a cell override are
  translated to ``vllm bench serve`` options. They cannot override model,
  endpoint, sequence lengths, concurrency, result paths, or harness-owned
  percentile reporting.

.. _vllm-sweep:

Sweep
=====

The sweep is an explicit list of canonical cells, not a cartesian product.
``sweeps`` defines optional overrides and ``runs`` selects the cells to execute.

.. code:: json

    {
      "sweeps": {
        "ISL=1000,OSL=1000,TP=8,PP=2,CONC=16": { "num_prompts": 50 },
        "ISL=1000,OSL=1000,TP=8,PP=2,CONC=32": {}
      },
      "runs": [
        "ISL=1000,OSL=1000,TP=8,PP=2,CONC=16",
        "ISL=1000,OSL=1000,TP=8,PP=2,CONC=32"
      ]
    }

.. list-table::
   :widths: 3 7
   :header-rows: 1

   * - Key
     - Description
   * - ``sweeps.<cell>``
     - Per-cell benchmark override object.
   * - ``runs[]``
     - Canonical key declared in ``sweeps``.

An undeclared, malformed, duplicated, or TP/PP-inconsistent cell key is a
load-time error.

Cell keys
---------

Each run is one **cell**, identified by a canonical key used to look up thresholds:

.. code:: text

  ISL=<isl>,OSL=<osl>,TP=<tp>,PP=<pp>,CONC=<conc>

``PP=`` is always present, including single-node and Ray runs with ``PP=1``.
The host count is placement information, not a threshold dimension.
Examples::

  ISL=1000,OSL=1000,TP=8,PP=1,CONC=16
  ISL=1000,OSL=1000,TP=8,PP=2,CONC=16

Server reuse
------------

Cells with identical server arguments share a server identity, so the suite
reuses the running server instead of stopping it, restarting, and reloading
weights. The derived ``--max-model-len`` is part of that identity: different
derived values force a restart, while cells with equal derived values reuse the
server. For example, with zero range ratio and prefix, 1024/8192 and 8192/1024
both derive ``9224`` and can share. An explicit
non-null ``server_params.max_model_len`` can also allow different ISL/OSL
cells to share. Explicit null removes the max-model-len option from the server
identity entirely, so different ISL/OSL cells share when their other server
arguments match. Concurrency remains client-only, so ordering runs with
concurrency varying fastest makes a sweep substantially quicker.

.. _vllm-thresholds:

Thresholds
==========

Thresholds are keyed by canonical cell, then by a bare metric name:

.. code:: json

    {
      "ISL=1000,OSL=1000,TP=8,PP=1,CONC=16": {
        "output_throughput": {"kind": "min", "value": 4000},
        "mean_ttft_ms": {"kind": "max", "value": 500},
        "failed": {"kind": "max", "value": 0}
      }
    }

A cell may contain any subset of the registry, including an empty object. When
``enforce_thresholds`` is true, every selected run must have a threshold cell,
but only specs present in that cell create verification subtests. When it is
false, produced and configured values remain ``record`` rows and no threshold
subtests run.

Sweep specs are strict at load time regardless of enforcement:

- A spec contains exactly ``kind`` and ``value``.
- ``kind`` must equal the registry direction, exactly ``min`` or ``max``.
- ``value`` must be a finite JSON number. Booleans, strings, null, arrays,
  objects, NaN, and infinity are rejected.
- Cell-level keys beginning with ``_comment`` or ``_example`` are metadata and
  are ignored by metric validation; all other keys must name a registered
  metric.
- Prefixed names (``client.*``, ``gpu.*``, ``prom.*``), unknown names, legacy
  kinds (including ``min_tok_s`` and ``max_ms``), references, tolerances,
  units, ``info``, and extra fields are rejected.

``min`` fails below its value and passes at or above it. ``max`` fails above
its value and passes at or below it. A gated actual that is missing, null,
boolean, string, collection, NaN, or infinity fails for every datasource.

The optional top-level ``accuracy`` block remains task-qualified and is exempt
from these vLLM sweep-name and spec rules.

.. _vllm-threshold-discovery:

Threshold file discovery
------------------------

Set ``threshold_json`` to the threshold file. A relative path resolves against
the configuration file's directory. The packaged 28 files contain two cells
each and intentionally retain the previous 32-metric assertion subset in
canonical registry order. All 56 registry metrics remain supported, and every
finite produced value is reported whether or not it has a threshold spec. The
packaged zero values are uncalibrated placeholders, and every paired config
keeps enforcement disabled. Users may add any other registered metric or
remove any packaged entry.

Metrics
=======

vLLM uses one ordered 56-entry bare-name registry. The metric name owns its
raw source or derivation inputs, display unit, datasource, display category,
and exact threshold direction. ``median_*`` and ``p50_*`` remain separate.

Run and health (7)
------------------

``max_concurrency``, ``max_concurrent_requests``, ``num_prompts``,
``completed``, ``failed``, ``success_rate``, ``duration``.

Throughput and totals (10)
--------------------------

``request_throughput``, ``goodput``, ``output_throughput``,
``total_token_throughput``, ``per_gpu_throughput``,
``decode_throughput_p50``, ``max_output_tokens_per_s``, ``rtfx``,
``total_input_tokens``, ``total_output_tokens``.

TTFT (8)
--------

``mean_ttft_ms``, ``median_ttft_ms``, ``std_ttft_ms``, ``p50_ttft_ms``,
``p90_ttft_ms``, ``p95_ttft_ms``, ``p99_ttft_ms``,
``normalized_ttft_ms_per_tok``.

TPOT (7)
--------

``mean_tpot_ms``, ``median_tpot_ms``, ``std_tpot_ms``, ``p50_tpot_ms``,
``p90_tpot_ms``, ``p95_tpot_ms``, ``p99_tpot_ms``.

ITL (8)
-------

``mean_itl_ms``, ``median_itl_ms``, ``std_itl_ms``, ``p50_itl_ms``,
``p90_itl_ms``, ``p95_itl_ms``, ``p99_itl_ms``, ``decode_latency_ratio``.

End-to-end latency (7)
----------------------

``mean_e2el_ms``, ``median_e2el_ms``, ``std_e2el_ms``, ``p50_e2el_ms``,
``p90_e2el_ms``, ``p95_e2el_ms``, ``p99_e2el_ms``.

GPU (5)
-------

``peak_gpu_memory_mb``, ``model_load_memory_mb``, ``model_load_s``,
``gpu_bandwidth_util_pct``, ``gpu_compute_util_pct``. Load time and memory are
captured once after successful readiness and reused for cells sharing that
server. Load time remains available whenever the elapsed measurement is finite.
Load memory requires finite pre/post VRAM snapshots; missing snapshots do not
disable server reuse or the elapsed measurement. A real zero memory delta
remains zero.

Prometheus (4)
--------------

``queue_time_p50_ms``, ``queue_time_p95_ms``, ``prefill_time_p50_ms``, and
``prefill_time_p95_ms``. CVS diffs before/after histogram scrapes so reused
server counters remain isolated to one cell.

Directions
----------

The following are ``min``: ``max_concurrency``,
``max_concurrent_requests``, ``num_prompts``, ``completed``, ``success_rate``,
``request_throughput``, ``goodput``, every token throughput/total metric,
``rtfx``, ``gpu_bandwidth_util_pct``, and ``gpu_compute_util_pct``. Every other
registry metric is ``max``.

Projection and derivation
-------------------------

The vLLM result projector accepts only finite built-in integer and float values;
booleans are not numeric. Known metadata is ignored: ``date``,
``endpoint_type``, ``backend``, ``label``, ``model_id``, ``tokenizer_id``,
``burstiness``, and ``request_rate`` (including a finite request rate). Any
unknown finite top-level numeric field fails parsing and names the artifact.
The raw ``request_goodput`` field maps only to ``goodput``.

Derived metrics are emitted only when their result is finite:

.. code:: text

  per_gpu_throughput          = total_token_throughput / (tp * pp)
  normalized_ttft_ms_per_tok  = mean_ttft_ms / isl
  decode_latency_ratio        = p99_itl_ms / p50_itl_ms
  decode_throughput_p50       = 1000 / median_tpot_ms
  success_rate                = completed / (completed + failed)

Reporting and compatibility
---------------------------

``test_verify_cell_metrics`` remains one parent per cell. It computes all host
rows before emitting one subtest for each present, enforced spec, so one failure
does not hide sibling verdicts. Finite produced values without a spec are
record-only HTML rows. The parent also emits one compact JUnit property with
``actuals_by_host`` and metric contract ``{"id":"vllm-bare","version":1}``.

Run Deck tables, charts, and highlights use selected registry metrics rather
than all 56 columns. Historical namespaced vLLM Run Deck artifacts do not match
the contract: previous-run and manually selected viewer baselines display an
explicit incompatibility and suppress performance comparisons. A compatible
task-qualified accuracy section in the same baseline remains independently
eligible for accuracy comparison.

Results table
-------------

The summary table emits seven fixed columns — Model, GPU, ISL, OSL, Policy,
Conc, Host — followed by Req/s, Total tok/s, Mean TTFT, P95 TTFT, Mean TPOT,
P95 TPOT, P99 ITL, and Goodput.

.. _vllm-accuracy:

Accuracy tests
==============

Accuracy evaluation runs `lm-evaluation-harness <https://github.com/EleutherAI/lm-evaluation-harness>`_ against the live server after the performance sweep. Task **selection** lives in the configuration file; **gating values** live in the threshold file.

.. code:: json

    {
      "accuracy": {
        "tasks": [
          {
            "id": "gsm8k_strict",
            "tasks": "gsm8k",
            "backend": "vllm",
            "lm_eval_model": "local-completions",
            "num_fewshot": 5,
            "batch_size": "auto",
            "limit": 100,
            "num_concurrent": 8,
            "exec_timeout_sec": 7200,
            "extra_model_args": "tokenizer_backend=huggingface"
          }
        ]
      }
    }

.. list-table::
   :widths: 3 2 5
   :header-rows: 1

   * - Key
     - Default
     - Description
   * - ``id``
     - none
     - Unique label for this entry; duplicates are rejected.
   * - ``tasks`` (or legacy ``task``)
     - none
     - lm-eval task name or task list.
   * - ``lm_eval_model``
     - endpoint-derived
     - ``local-completions`` or ``local-chat-completions``
   * - ``num_fewshot``
     - lm-eval default
     - Few-shot example count, if explicitly set.
   * - ``num_concurrent``
     - ``8``
     - Concurrent requests.
   * - ``apply_chat_template``
     - ``false``
     - Enables or names the chat template.
   * - ``metadata``
     - ``{}``
     - Passed through to lm-eval.
   * - ``include_path``
     - ``""``
     - Directory of custom task definitions.
   * - ``gen_kwargs``
     - ``{}``
     - Generation arguments.

``lm_eval_model`` selects the API surface. If omitted, CVS derives it from
``apply_chat_template``:

.. list-table::
   :widths: 2 3 4
   :header-rows: 1

   * - Value
     - lm-eval model
     - Endpoint
   * - ``false``
     - ``local-completions``
     - ``/v1/completions``
   * - ``true``
     - ``local-chat-completions``
     - ``/v1/chat/completions``

CVS pins runtime installation to ``lm-eval[api,math]==0.4.12``. The shared
accuracy schema also exposes that release's evaluation controls, including
``batch_size``, ``max_batch_size``, ``device``, ``limit``, ``samples``,
``use_cache``, ``cache_requests``, ``check_integrity``,
``system_instruction``, ``fewshot_as_multiturn``, ``predict_only``, ``seed``,
``trust_remote_code``, ``confirm_run_unsafe_code``, ``metadata``, and
``gen_kwargs``. CVS owns the endpoint, model path, output path, and sample
logging. Results land under ``<log_dir>/accuracy``.

Accuracy metric keys
--------------------

Accuracy metrics are keyed ``<lm_task_name>.<metric_key>``, with any comma in the name replaced by a double underscore. For example, gsm8k's ``exact_match,strict-match`` becomes:

.. code:: text

  gsm8k.exact_match__strict-match

Gate them in the threshold file's top-level ``accuracy`` block, keyed by the task ``id``:

.. code:: json

    {
      "accuracy": {
        "gsm8k_strict": {
          "gsm8k.exact_match__strict-match": { "kind": "min", "value": 0.75 }
        }
      }
    }

.. note::

  An accuracy failure does not mark the shared lifecycle as failed, so the remaining stages — the results table and teardown — still run normally.

Troubleshooting
===============

.. list-table::
   :widths: 5 5
   :header-rows: 1

   * - Message
     - Cause and fix
   * - Distributed execution requires ``pipeline_parallel_size > 1``
     - Multi-host ``vllm_distributed`` on the mp backend needs pipeline parallelism. Either raise ``pipeline_parallel_size``, or set ``distributed-executor-backend`` to ``"ray"``.
   * - ``vllm_single requires pipeline_parallel_size=1``
     - Use ``vllm_distributed`` when the config requires pipeline parallelism.
   * - ``vllm_distributed requires container.env.NCCL_SOCKET_IFNAME``
     - Set all three socket-interface variables under ``container.env``.
   * - ``Container image not specified in config``
     - ``container.image`` is empty. Note that a variant ``container`` block with no ``image`` overwrites the cluster file's value.
   * - ``runs must be a nonempty explicit list``
     - ``runs`` is missing or empty. List at least one canonical cell key from ``sweeps``.
   * - ``runs contains duplicate cells``
     - The same cell key appears twice in ``runs``.
   * - ``runs reference unknown sweeps``
     - A ``runs`` entry is not a key in ``sweeps``.
   * - ``run cell must be canonical ISL=<n>,OSL=<n>,TP=<n>,PP=<n>,CONC=<n>``
     - A sweep or run key is malformed.
   * - ``conflicts with server_params tensor/pipeline parallel size``
     - The TP or PP in a cell key does not match ``server_params``.
   * - ``duplicate task id(s)``
     - Two ``accuracy.tasks`` entries share an ``id``.
   * - ``unknown vLLM threshold metric '<metric>'``
     - The cell key does not begin with ``_`` and does not name one of the 56 registry metrics.
   * - ``<metric> threshold kind must be '<direction>', got '<kind>'``
     - ``kind`` does not match the registry direction. Use the required ``min`` or ``max`` value shown in the message.
   * - ``<metric>: actual must be a finite built-in int or float, got <value>``
     - The gated datasource did not produce a valid value. Inspect the benchmark artifact, GPU telemetry, or server metrics for that cell; percentile collection is harness-owned.
   * - ``NotImplementedError: model.remote=1``
     - Remote model download is unimplemented. Pre-stage weights and set ``remote: 0``.
   * - ``ValueError: too many values to unpack``
     - ``env`` was placed under ``runtime.args``. Move it to the ``container`` top level.
   * - Extra-key validation error
     - A misspelled key. Every block except ``container`` forbids unknown keys.

Related resources
=================

- :doc:`/how-to/test-suites/inference/vllm` — step-by-step first run
- :doc:`/reference/cluster/cluster-file` — cluster file and orchestrator backends
- :doc:`/how-to/run-with-containers` — container backend walkthrough
- :doc:`/how-to/test-suites/index` — running other CVS suites
