vLLM inference benchmark configuration file for Cluster Validation Suite (CVS)#

2026-09-24

24 min read time

Applies to Linux

The vLLM suites benchmark LLM serving throughput, latency, and accuracy on AMD Instinct GPUs. vllm_single runs on the first cluster host and ignores additional hosts. vllm_distributed uses every host in the cluster file, with one-host fallback when only a single host is present. Packaged distributed recipes and thresholds are calibrated for two hosts; retune them before treating other sizes as pass/fail.

For a mapping-style node_dict, “first cluster host” means the first JSON key in insertion order. vllm_single scopes execution to that host and rewrites head_node_dict.mgmt_ip to match, even if the original head names another node. Put the intended single-node target first in node_dict.

Run it with:

cvs run vllm_single --cluster_file <cluster.json> --config_file <config.json>

For a step-by-step walkthrough of a first run, see Run vLLM LLM inference tests with CVS. This page is the schema and metric reference.

Lifecycle#

Each stage of the run is an independent test, so every stage becomes its own timed, pass/fail row in the HTML report. The suite pins this order explicitly rather than relying on definition order:

Order

Test

Purpose

0

test_launch_container

Pull/load the image and start the container on every node.

1

test_setup_sshd

Always skipped for vLLM (see note below).

2

test_discover_topology

Resolve IB HCA devices; no-op for effective single-node execution.

3

test_model_fetch

Stage model weights.

4

test_openai_compatible_smoke

Short-lived server; verifies the OpenAI-compatible API answers.

5

test_vllm_inference

Run one benchmark cell (parametrized per sweep run).

6

test_verify_cell_metrics

One verification parent per cell; configured threshold gates are listed as subtests.

7

test_accuracy_eval

lm-eval accuracy tasks, if any are configured.

8

test_print_results_table

Console + report summary table.

9

test_teardown

Stop the server and tear down the container.

Note

test_setup_sshd always skips in this suite. vLLM uses --distributed-executor-backend mp with NCCL over the host network, so no inter-container sshd is needed. A skipped row here is expected, not a problem.

Note

vLLM suite execution is serial and single-pass. Do not use xdist workers or pytest-repeat counts above one: the verification phase consumes results collected earlier in the same pytest process.

If a stage fails, later stages are skipped rather than cascading into confusing downstream errors. The container is still torn down by a leak-guard even when a mid-sweep test fails.

Configuration file structure#

A vLLM configuration file has these top-level keys:

Key

Required

Description

enforce_thresholds

no (default true)

When false, metrics record without requiring calibrated threshold cells.

threshold_json

yes

Explicit path to the threshold file. See Threshold file discovery.

container

yes

Container/Docker settings. See Container and Docker configuration.

paths

yes

Filesystem locations. See Paths.

server_params

yes

Harness-owned server fields plus snake-case vllm serve options.

benchmark_params

no

Benchmark defaults plus snake-case vllm bench serve options.

sweeps

yes

Canonical run-cell keys mapped to benchmark overrides.

runs

yes

Nonempty ordered list of the sweep cells to execute.

thresholds

no

Per-cell pass/fail specs. See Thresholds.

accuracy

no

lm-eval task selection. See Accuracy tests.

Important

Top-level and structural blocks forbid unknown keys. server_params, benchmark_params, and individual sweep overrides deliberately accept arbitrary snake-case vLLM option names. CVS converts them to kebab-case CLI flags. null omits a flag, true emits a bare flag, scalars emit one value, lists emit one flag followed by values, and mappings emit compact JSON. Use an option’s negative form instead of false.

Placeholder substitution#

Values are resolved in three passes, so later forms can reference earlier ones:

  1. Cluster placeholders — {user-id} resolves to the current username.

  2. Self-reference within paths — {shared_fs} expands to the already-resolved paths.shared_fs.

  3. Cross-block — {paths.models_dir} expands anywhere else in the file, such as in a volume mount.

{
  "paths": {
    "shared_fs": "/mnt/dtni/{user-id}",
    "models_dir": "{shared_fs}/models",
    "log_dir": "{shared_fs}/LOGS",
    "hf_token_file": "{shared_fs}/.cache/huggingface/token"
  },
  "container": {
    "runtime": {
      "args": {
        "volumes": ["{paths.models_dir}:/models"]
      }
    }
  }
}

Execution backends#

Four different things in this stack are called a “backend”. They are unrelated, and confusing them is the most common configuration mistake.

Setting

Values

What it selects

benchmark_params.backend

"vllm" (default)

The client backend passed to vllm bench serve --backend. Nothing to do with distribution.

server_params.distributed_executor_backend

"mp" (default), "ray"

How vLLM distributes the model across nodes. This is the multinode setting.

container.runtime.name

"docker"

The container runtime.

Cluster file orchestrator

"baremetal", "container"

Whether CVS runs commands on the host or inside a container. See Cluster Validation Suite (CVS) cluster file: configuration and backend selection.

Warning

enroot is registered but not implemented. Every method is a stub that returns failure, so a run with runtime.name: "enroot" fails at test_launch_container. Podman is not supported. Use docker.

Distributed executor: mp and ray#

Multinode runs support two executor backends. mp is the default and requires no configuration key at all.

mp (default). Used whenever server_params.distributed_executor_backend is absent. The suite injects the full distributed block into each rank’s vllm serve command:

vllm serve <model> --tensor-parallel-size <tp> --port <port> \
  --node-rank <rank> --master-addr <addr> --master-port <port> \
  --nnodes <n> --pipeline-parallel-size <pp> \
  --distributed-executor-backend mp

Every rank above 0 additionally gets --headless. This path requires pipeline parallelism (pipeline_parallel_size greater than 1).

ray (opt-in). Selected by setting server_params.distributed_executor_backend to the exact lowercase string "ray". Other values are configuration errors. Ray takes a completely different route:

  1. Bootstrap the cluster head: ray start --head --port=<master_port>

  2. Bootstrap each worker: ray start --address=<cluster-head>:<dist_init_port>

  3. Launch vllm serve on the head node only — workers run no serve process

  4. On teardown, broadcast ray stop after the process kill

Under ray, none of the mp distributed flags are emitted. --pipeline-parallel-size is added only when server_params.pipeline_parallel_size is greater than 1.

Note

Ray does not enable multinode — it relaxes the pipeline-parallelism requirement. With ray, pipeline_parallel_size of 1 is legal and is the expected configuration for pure tensor-parallel multinode serving. With mp, pipeline parallelism is mandatory.

Because only the head node serves under ray, worker ranks produce no per-rank server log. That is expected.

Topology validation rules#

These rules are enforced when the configuration file loads, before anything starts:

Condition

Rule

two or more cluster hosts, backend is not ray

pipeline_parallel_size must be greater than 1

two or more cluster hosts, backend is ray

pipeline_parallel_size of 1 is valid

pipeline_parallel_size > 1

The distributed suite requires more than one cluster host

two or more cluster hosts, either backend

container.env.NCCL_SOCKET_IFNAME is required

The corresponding error messages are:

multi-host distributed execution requires pipeline_parallel_size > 1 unless using ray
pipeline_parallel_size > 1 requires a multi-host distributed suite
vllm_distributed requires container.env.NCCL_SOCKET_IFNAME on multi-host clusters

Multinode prerequisites#

Beyond the validation rules, a multinode run needs:

  • server_params.dist_init_port — default 29501; CVS derives the head address from the cluster.

  • container.env.NCCL_IB_HCA — the comma-separated RDMA HCA names available on every node. The packaged MI3xx configurations set rdma0 through rdma7.

  • container.env.NCCL_SOCKET_IFNAME — the Linux netdev associated with the selected RNICs.

  • container.env.GLOO_SOCKET_IFNAME and TP_SOCKET_IFNAME — generally the frontend/control-plane interface.

  • container.env.NCCL_IB_GID_INDEX — the index for the intended RoCE/IB fabric. Use show_gids inside the container and choose an entry available on every selected HCA and node. If that command is unavailable, inspect ibv_devinfo -v and /sys/class/infiniband/<hca>/ports/<port>/gid_attrs/.

Container and Docker configuration#

The container block controls image selection, lifetime, and the docker run flags.

{
  "container": {
    "lifetime": "per_run",
    "name": "vllm_perf_inference_rocm",
    "image": "rocm/vllm:latest",
    "runtime": {
      "name": "docker",
      "args": {
        "network": "host",
        "ipc": "host",
        "privileged": true,
        "volumes": [
          "/home/{user-id}:/home/{user-id}",
          "{paths.models_dir}:/models"
        ]
      }
    }
  }
}

Container block keys#

The following keys are accepted inside the container block.

Key

Default

Description

image

none

Container image. Required — launch fails with Container image not specified in config.

name

<user>_<sanitized-image>

Container name.

lifetime

"per_run"

One of no_launch, per_run, persistent.

runtime.name

"docker"

Container runtime.

runtime.args

{}

Docker flags; see the table below.

env

{}

Container-level environment variables. Top level, not under runtime.args.

image_tar

absent

Path on each host to a saved image tar to docker load instead of pulling. Top level.

Warning

Put env at the container top level. Placing it under runtime.args crashes container launch: the code iterates that value as a sequence of pairs, which raises ValueError: too many values to unpack for any key longer than two characters.

Runtime arguments#

All keys under runtime.args are optional. List-valued keys append to the defaults; scalar keys override them. There is no way to remove a default device or capability.

Key

Merge

Default

Emitted flag

volumes

append

/home/$USER/.ssh:/host_ssh (always added)

-v <host>:<ctr>[:ro]

devices

append

/dev/kfd, /dev/dri, /dev/infiniband

--device <path>

cap_add

append

SYS_PTRACE, IPC_LOCK, SYS_ADMIN

--cap-add <cap>

security_opt

append

seccomp=unconfined, apparmor=unconfined

--security-opt <opt>

group_add

append

video

--group-add <group>

ulimit

append

memlock=-1

--ulimit <limit>

network

override

host

--network <mode>

ipc

override

host

--ipc <mode>

privileged

override

true

--privileged

registry

n/a

none

Triggers docker login; see below

The assembled command is:

docker run -d --name <name> <args> <image> sleep infinity

The container is a long-lived sidecar; every workload command runs through docker exec inside it. InfiniBand devices are additionally passed through by per-host shell expansion at launch time, so each node mounts the devices it actually has.

Note

--gpus is deliberately never emitted — GPU access on AMD hardware comes from the /dev/kfd and /dev/dri device mounts plus the video group.

shm_size is not supported on this path. Setting runtime.args.shm_size is silently ignored; --shm-size is never emitted.

Container lifetime#

The lifetime key controls when the container is started and stopped relative to the test lifecycle.

lifetime

Setup behavior

Teardown behavior

no_launch

Verifies a container of that name is already running; never starts one.

No-op

per_run

Force-removes any stale container of the same name, then launches.

docker rm -f

persistent

Attaches if running on all hosts; cold-starts if absent on all hosts; refuses on partial or failed probe.

No-op

Tip

With persistent, always pin container.name explicitly. The default name is derived from the image, so bumping an image tag silently abandons the old container and starts a new one.

Registry authentication#

Set runtime.args.registry to log in before pulling:

{
  "registry": {
    "username": "myuser",
    "password_file": "/path/on/each/host/to/token",
    "server": "registry.example.com"
  }
}

username and password_file are required; server defaults to Docker Hub. The password is read from a file path on each remote host — there is no inline password or token key, and the login is kept out of the logs. Login is skipped entirely when image_tar is set, since a tar load never pulls.

Image resolution order at launch:

  1. If image_tar is set and the image is absent, docker load it.

  2. Otherwise, if registry is set, log in.

  3. Check whether the image exists on all hosts.

  4. If not, docker pull it, with no retry or fallback.

Note

The image-exists check matches Repository:Tag exactly, so an image referenced without a tag or by digest never matches and is pulled on every run.

Cluster file merge#

The variant’s container block is deep-merged onto the cluster file’s block: dictionaries merge key-wise, while scalars and lists are replaced. Cluster-set values survive unless the variant sets the same key.

Warning

A container block in the variant always contributes lifetime, name, and image — including empty defaults. If your variant defines container but omits image, it overwrites a cluster-file image with an empty string and the launch fails. Set image in whichever file defines the block.

Paths#

All four keys are required.

Key

Description

shared_fs

Root of the shared filesystem, typically the anchor other paths reference.

models_dir

Model weight cache; exported into the server as HF_HUB_CACHE.

log_dir

Root for run artifacts.

hf_token_file

Path to a file containing the Hugging Face token.

If hf_token_file does not exist, the run continues with an empty token. vLLM configs always serve a pre-staged model mounted under paths.models_dir.

Per-cell artifacts land in:

<log_dir>/vllm/out-node<rank>/isl<isl>_osl<osl>_conc<conc>/
  vllm_serve_server.log
  client.log
  results

Model#

Key

Default

Description

server_params.model

none

Local path, for example /models/Llama-3.1-70B-Instruct-FP8-KV

Server role#

server_params controls the vllm serve process. Harness-owned fields are model, tensor_parallel_size, pipeline_parallel_size, port, dist_init_port, polling controls, and distributed_executor_backend. Every other snake-case key is passed through to vllm serve.

Key

Default

Description

model

none

Local model path supplied as the positional vllm serve argument.

tensor_parallel_size

none

Tensor-parallel degree.

pipeline_parallel_size

1

Pipeline-parallel degree.

port

8888

OpenAI-compatible server port.

How server_params are flattened#

JSON value

Emitted

Example

Scalar

--flag value

"kv_cache_dtype": "fp8" → --kv-cache-dtype fp8

true

--flag (bare)

"enforce_eager": true → --enforce-eager

false

rejected

Omit the setting or use vLLM’s explicit negative option.

List

One flag followed by its values.

"x": ["a","b"] → --x a b

When server_params.max_model_len is absent, CVS derives one for each benchmark cell from its effective parameters after sweep overrides:

ceil((ISL + OSL) * (1 + random_range_ratio)) + random_prefix_len + 8

An explicit non-null server_params.max_model_len takes precedence and emits exactly one --max-model-len flag. An explicit null value is an intentional opt-out: the generic option serializer emits no flag, and CVS suppresses the fallback so vLLM uses the model or image default.

Environment variables: two mechanisms#

These are separate and are frequently confused.

container.env

generated per-command environment

Applied by

docker run -e

A sourced shell script inside the container after HCA discovery.

Scope

Every command in the container, for its whole lifetime.

Hugging Face path/token variables and legacy network fallbacks.

Changing it

Requires recreating the container.

Takes effect on the next command.

Defaults

GPUS=8, MULTINODE=true

No network overrides unless a legacy top-level field is set.

The generated per-command environment always exports:

export HF_TOKEN=<token>
export HF_HUB_CACHE=<paths.models_dir>

Packaged configurations set the HCA, socket-interface, GID, and NCCL debug settings in container.env so every command inherits them and topology discovery does not overwrite them. The legacy top-level ib_hca_devices and ib_netdev fields remain fallbacks for external configurations that omit their corresponding container environment variables. Put other static ROCm, NCCL, and vLLM exports in container.env.

The packaged MI3xx catalog’s AITER settings are image- and model-scoped. For the image based on vLLM commit 4bdc8a788:

  • DeepSeek V4 Flash, DeepSeek V4 Pro, GLM 5.1, and GLM 5.2 set VLLM_ROCM_USE_AITER=1, VLLM_ROCM_USE_AITER_MHA=0, and GPU_ARCHS=gfx942.

  • Kimi K2.5 sets VLLM_ROCM_USE_AITER=1 and disables VLLM_ROCM_USE_AITER_MHA, VLLM_ROCM_USE_AITER_FP4BMM, VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS, and VLLM_ROCM_USE_AITER_MLA.

  • Other packaged model families carry no AITER overrides.

Do not add VLLM_USE_AITER_UNIFIED_ATTENTION or VLLM_ROCM_USE_AITER_FUSED_MOE_A16W4: this image does not register them, so they are warning-only and ignored. Do not substitute VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION or other attention-backend variables without validating both registration and call sites in the exact image.

Benchmark parameters#

benchmark_params holds client defaults. Per-cell entries in sweeps override these values. CVS owns endpoint construction, result paths, and percentile reporting.

Key

Default

Description

backend

"vllm"

Client backend for vllm bench serve.

base_url

"http://0.0.0.0"

Server base URL.

num_prompts

3200

Total prompts per cell.

dataset_name

"random"

Dataset for the load generator.

num_prompts

"3200"

Total prompts per cell.

burstiness

"1.0"

1.0 is a uniform arrival process; lower is burstier.

seed

"0"

Random seed.

request_rate

"inf"

Arrival rate; inf sends as fast as concurrency allows.

random_range_ratio

"0.0"

Length jitter around ISL/OSL; also feeds the derived max-model-len.

random_prefix_len

"0"

Shared prefix length.

tokenizer_mode

"auto"

Tokenizer mode.

client_poll_iterations

20

Client completion polls before giving up.

Tip

Arbitrary snake-case keys under benchmark_params and a cell override are translated to vllm bench serve options. They cannot override model, endpoint, sequence lengths, concurrency, result paths, or harness-owned percentile reporting.

Sweep#

The sweep is an explicit list of canonical cells, not a cartesian product. sweeps defines optional overrides and runs selects the cells to execute.

{
  "sweeps": {
    "ISL=1000,OSL=1000,TP=8,PP=2,CONC=16": { "num_prompts": 50 },
    "ISL=1000,OSL=1000,TP=8,PP=2,CONC=32": {}
  },
  "runs": [
    "ISL=1000,OSL=1000,TP=8,PP=2,CONC=16",
    "ISL=1000,OSL=1000,TP=8,PP=2,CONC=32"
  ]
}

Key

Description

sweeps.<cell>

Per-cell benchmark override object.

runs[]

Canonical key declared in sweeps.

An undeclared, malformed, duplicated, or TP/PP-inconsistent cell key is a load-time error.

Cell keys#

Each run is one cell, identified by a canonical key used to look up thresholds:

ISL=<isl>,OSL=<osl>,TP=<tp>,PP=<pp>,CONC=<conc>

PP= is always present, including single-node and Ray runs with PP=1. The host count is placement information, not a threshold dimension. Examples:

ISL=1000,OSL=1000,TP=8,PP=1,CONC=16
ISL=1000,OSL=1000,TP=8,PP=2,CONC=16

Server reuse#

Cells with identical server arguments share a server identity, so the suite reuses the running server instead of stopping it, restarting, and reloading weights. The derived --max-model-len is part of that identity: different derived values force a restart, while cells with equal derived values reuse the server. For example, with zero range ratio and prefix, 1024/8192 and 8192/1024 both derive 9224 and can share. An explicit non-null server_params.max_model_len can also allow different ISL/OSL cells to share. Explicit null removes the max-model-len option from the server identity entirely, so different ISL/OSL cells share when their other server arguments match. Concurrency remains client-only, so ordering runs with concurrency varying fastest makes a sweep substantially quicker.

Thresholds#

Thresholds are keyed by canonical cell, then by a bare metric name:

{
  "ISL=1000,OSL=1000,TP=8,PP=1,CONC=16": {
    "output_throughput": {"kind": "min", "value": 4000},
    "mean_ttft_ms": {"kind": "max", "value": 500},
    "failed": {"kind": "max", "value": 0}
  }
}

A cell may contain any subset of the registry, including an empty object. When enforce_thresholds is true, every selected run must have a threshold cell, but only specs present in that cell create verification subtests. When it is false, produced and configured values remain record rows and no threshold subtests run.

Sweep specs are strict at load time regardless of enforcement:

  • A spec contains exactly kind and value.

  • kind must equal the registry direction, exactly min or max.

  • value must be a finite JSON number. Booleans, strings, null, arrays, objects, NaN, and infinity are rejected.

  • Cell-level keys beginning with _comment or _example are metadata and are ignored by metric validation; all other keys must name a registered metric.

  • Prefixed names (client.*, gpu.*, prom.*), unknown names, legacy kinds (including min_tok_s and max_ms), references, tolerances, units, info, and extra fields are rejected.

min fails below its value and passes at or above it. max fails above its value and passes at or below it. A gated actual that is missing, null, boolean, string, collection, NaN, or infinity fails for every datasource.

The optional top-level accuracy block remains task-qualified and is exempt from these vLLM sweep-name and spec rules.

Threshold file discovery#

Set threshold_json to the threshold file. A relative path resolves against the configuration file’s directory. The packaged 28 files contain two cells each and intentionally retain the previous 32-metric assertion subset in canonical registry order. All 56 registry metrics remain supported, and every finite produced value is reported whether or not it has a threshold spec. The packaged zero values are uncalibrated placeholders, and every paired config keeps enforcement disabled. Users may add any other registered metric or remove any packaged entry.

Metrics#

vLLM uses one ordered 56-entry bare-name registry. The metric name owns its raw source or derivation inputs, display unit, datasource, display category, and exact threshold direction. median_* and p50_* remain separate.

Run and health (7)#

max_concurrency, max_concurrent_requests, num_prompts, completed, failed, success_rate, duration.

Throughput and totals (10)#

request_throughput, goodput, output_throughput, total_token_throughput, per_gpu_throughput, decode_throughput_p50, max_output_tokens_per_s, rtfx, total_input_tokens, total_output_tokens.

TTFT (8)#

mean_ttft_ms, median_ttft_ms, std_ttft_ms, p50_ttft_ms, p90_ttft_ms, p95_ttft_ms, p99_ttft_ms, normalized_ttft_ms_per_tok.

TPOT (7)#

mean_tpot_ms, median_tpot_ms, std_tpot_ms, p50_tpot_ms, p90_tpot_ms, p95_tpot_ms, p99_tpot_ms.

ITL (8)#

mean_itl_ms, median_itl_ms, std_itl_ms, p50_itl_ms, p90_itl_ms, p95_itl_ms, p99_itl_ms, decode_latency_ratio.

End-to-end latency (7)#

mean_e2el_ms, median_e2el_ms, std_e2el_ms, p50_e2el_ms, p90_e2el_ms, p95_e2el_ms, p99_e2el_ms.

GPU (5)#

peak_gpu_memory_mb, model_load_memory_mb, model_load_s, gpu_bandwidth_util_pct, gpu_compute_util_pct. Load time and memory are captured once after successful readiness and reused for cells sharing that server. Load time remains available whenever the elapsed measurement is finite. Load memory requires finite pre/post VRAM snapshots; missing snapshots do not disable server reuse or the elapsed measurement. A real zero memory delta remains zero.

Prometheus (4)#

queue_time_p50_ms, queue_time_p95_ms, prefill_time_p50_ms, and prefill_time_p95_ms. CVS diffs before/after histogram scrapes so reused server counters remain isolated to one cell.

Directions#

The following are min: max_concurrency, max_concurrent_requests, num_prompts, completed, success_rate, request_throughput, goodput, every token throughput/total metric, rtfx, gpu_bandwidth_util_pct, and gpu_compute_util_pct. Every other registry metric is max.

Projection and derivation#

The vLLM result projector accepts only finite built-in integer and float values; booleans are not numeric. Known metadata is ignored: date, endpoint_type, backend, label, model_id, tokenizer_id, burstiness, and request_rate (including a finite request rate). Any unknown finite top-level numeric field fails parsing and names the artifact. The raw request_goodput field maps only to goodput.

Derived metrics are emitted only when their result is finite:

per_gpu_throughput          = total_token_throughput / (tp * pp)
normalized_ttft_ms_per_tok  = mean_ttft_ms / isl
decode_latency_ratio        = p99_itl_ms / p50_itl_ms
decode_throughput_p50       = 1000 / median_tpot_ms
success_rate                = completed / (completed + failed)

Reporting and compatibility#

test_verify_cell_metrics remains one parent per cell. It computes all host rows before emitting one subtest for each present, enforced spec, so one failure does not hide sibling verdicts. Finite produced values without a spec are record-only HTML rows. The parent also emits one compact JUnit property with actuals_by_host and metric contract {"id":"vllm-bare","version":1}.

Run Deck tables, charts, and highlights use selected registry metrics rather than all 56 columns. Historical namespaced vLLM Run Deck artifacts do not match the contract: previous-run and manually selected viewer baselines display an explicit incompatibility and suppress performance comparisons. A compatible task-qualified accuracy section in the same baseline remains independently eligible for accuracy comparison.

Results table#

The summary table emits seven fixed columns — Model, GPU, ISL, OSL, Policy, Conc, Host — followed by Req/s, Total tok/s, Mean TTFT, P95 TTFT, Mean TPOT, P95 TPOT, P99 ITL, and Goodput.

Accuracy tests#

Accuracy evaluation runs lm-evaluation-harness against the live server after the performance sweep. Task selection lives in the configuration file; gating values live in the threshold file.

{
  "accuracy": {
    "tasks": [
      {
        "id": "gsm8k_strict",
        "tasks": "gsm8k",
        "backend": "vllm",
        "lm_eval_model": "local-completions",
        "num_fewshot": 5,
        "batch_size": "auto",
        "limit": 100,
        "num_concurrent": 8,
        "exec_timeout_sec": 7200,
        "extra_model_args": "tokenizer_backend=huggingface"
      }
    ]
  }
}

Key

Default

Description

id

none

Unique label for this entry; duplicates are rejected.

tasks (or legacy task)

none

lm-eval task name or task list.

lm_eval_model

endpoint-derived

local-completions or local-chat-completions

num_fewshot

lm-eval default

Few-shot example count, if explicitly set.

num_concurrent

8

Concurrent requests.

apply_chat_template

false

Enables or names the chat template.

metadata

{}

Passed through to lm-eval.

include_path

""

Directory of custom task definitions.

gen_kwargs

{}

Generation arguments.

lm_eval_model selects the API surface. If omitted, CVS derives it from apply_chat_template:

Value

lm-eval model

Endpoint

false

local-completions

/v1/completions

true

local-chat-completions

/v1/chat/completions

CVS pins runtime installation to lm-eval[api,math]==0.4.12. The shared accuracy schema also exposes that release’s evaluation controls, including batch_size, max_batch_size, device, limit, samples, use_cache, cache_requests, check_integrity, system_instruction, fewshot_as_multiturn, predict_only, seed, trust_remote_code, confirm_run_unsafe_code, metadata, and gen_kwargs. CVS owns the endpoint, model path, output path, and sample logging. Results land under <log_dir>/accuracy.

Accuracy metric keys#

Accuracy metrics are keyed <lm_task_name>.<metric_key>, with any comma in the name replaced by a double underscore. For example, gsm8k’s exact_match,strict-match becomes:

gsm8k.exact_match__strict-match

Gate them in the threshold file’s top-level accuracy block, keyed by the task id:

{
  "accuracy": {
    "gsm8k_strict": {
      "gsm8k.exact_match__strict-match": { "kind": "min", "value": 0.75 }
    }
  }
}

Note

An accuracy failure does not mark the shared lifecycle as failed, so the remaining stages — the results table and teardown — still run normally.

Troubleshooting#

Message

Cause and fix

Distributed execution requires pipeline_parallel_size > 1

Multi-host vllm_distributed on the mp backend needs pipeline parallelism. Either raise pipeline_parallel_size, or set distributed-executor-backend to "ray".

vllm_single requires pipeline_parallel_size=1

Use vllm_distributed when the config requires pipeline parallelism.

vllm_distributed requires container.env.NCCL_SOCKET_IFNAME

Set all three socket-interface variables under container.env.

Container image not specified in config

container.image is empty. Note that a variant container block with no image overwrites the cluster file’s value.

runs must be a nonempty explicit list

runs is missing or empty. List at least one canonical cell key from sweeps.

runs contains duplicate cells

The same cell key appears twice in runs.

runs reference unknown sweeps

A runs entry is not a key in sweeps.

run cell must be canonical ISL=<n>,OSL=<n>,TP=<n>,PP=<n>,CONC=<n>

A sweep or run key is malformed.

conflicts with server_params tensor/pipeline parallel size

The TP or PP in a cell key does not match server_params.

duplicate task id(s)

Two accuracy.tasks entries share an id.

unknown vLLM threshold metric '<metric>'

The cell key does not begin with _ and does not name one of the 56 registry metrics.

<metric> threshold kind must be '<direction>', got '<kind>'

kind does not match the registry direction. Use the required min or max value shown in the message.

<metric>: actual must be a finite built-in int or float, got <value>

The gated datasource did not produce a valid value. Inspect the benchmark artifact, GPU telemetry, or server metrics for that cell; percentile collection is harness-owned.

NotImplementedError: model.remote=1

Remote model download is unimplemented. Pre-stage weights and set remote: 0.

ValueError: too many values to unpack

env was placed under runtime.args. Move it to the container top level.

Extra-key validation error

A misspelled key. Every block except container forbids unknown keys.