vLLM inference benchmark configuration file for Cluster Validation Suite (CVS)#
2026-09-24
24 min read time
The vLLM suites benchmark LLM serving throughput, latency, and accuracy on AMD Instinct GPUs. vllm_single runs on the first cluster host and ignores additional hosts. vllm_distributed uses every host in the cluster file, with one-host fallback when only a single host is present. Packaged distributed recipes and thresholds are calibrated for two hosts; retune them before treating other sizes as pass/fail.
For a mapping-style node_dict, “first cluster host” means the first JSON key
in insertion order. vllm_single scopes execution to that host and rewrites
head_node_dict.mgmt_ip to match, even if the original head names another
node. Put the intended single-node target first in node_dict.
Run it with:
cvs run vllm_single --cluster_file <cluster.json> --config_file <config.json>
For a step-by-step walkthrough of a first run, see Run vLLM LLM inference tests with CVS. This page is the schema and metric reference.
Lifecycle#
Each stage of the run is an independent test, so every stage becomes its own timed, pass/fail row in the HTML report. The suite pins this order explicitly rather than relying on definition order:
Order |
Test |
Purpose |
|---|---|---|
0 |
|
Pull/load the image and start the container on every node. |
1 |
|
Always skipped for vLLM (see note below). |
2 |
|
Resolve IB HCA devices; no-op for effective single-node execution. |
3 |
|
Stage model weights. |
4 |
|
Short-lived server; verifies the OpenAI-compatible API answers. |
5 |
|
Run one benchmark cell (parametrized per sweep run). |
6 |
|
One verification parent per cell; configured threshold gates are listed as subtests. |
7 |
|
lm-eval accuracy tasks, if any are configured. |
8 |
|
Console + report summary table. |
9 |
|
Stop the server and tear down the container. |
Note
test_setup_sshd always skips in this suite. vLLM uses --distributed-executor-backend mp with NCCL over the host network, so no inter-container sshd is needed. A skipped row here is expected, not a problem.
Note
vLLM suite execution is serial and single-pass. Do not use xdist workers or
pytest-repeat counts above one: the verification phase consumes results
collected earlier in the same pytest process.
If a stage fails, later stages are skipped rather than cascading into confusing downstream errors. The container is still torn down by a leak-guard even when a mid-sweep test fails.
Configuration file structure#
A vLLM configuration file has these top-level keys:
Key |
Required |
Description |
|---|---|---|
|
no (default |
When |
|
yes |
Explicit path to the threshold file. See Threshold file discovery. |
|
yes |
Container/Docker settings. See Container and Docker configuration. |
|
yes |
Filesystem locations. See Paths. |
|
yes |
Harness-owned server fields plus snake-case |
|
no |
Benchmark defaults plus snake-case |
|
yes |
Canonical run-cell keys mapped to benchmark overrides. |
|
yes |
Nonempty ordered list of the sweep cells to execute. |
|
no |
Per-cell pass/fail specs. See Thresholds. |
|
no |
lm-eval task selection. See Accuracy tests. |
Important
Top-level and structural blocks forbid unknown keys. server_params,
benchmark_params, and individual sweep overrides deliberately accept
arbitrary snake-case vLLM option names. CVS converts them to kebab-case CLI
flags. null omits a flag, true emits a bare flag, scalars emit one
value, lists emit one flag followed by values, and mappings emit compact JSON.
Use an option’s negative form instead of false.
Placeholder substitution#
Values are resolved in three passes, so later forms can reference earlier ones:
Cluster placeholders —
{user-id}resolves to the current username.Self-reference within
paths—{shared_fs}expands to the already-resolvedpaths.shared_fs.Cross-block —
{paths.models_dir}expands anywhere else in the file, such as in a volume mount.
{
"paths": {
"shared_fs": "/mnt/dtni/{user-id}",
"models_dir": "{shared_fs}/models",
"log_dir": "{shared_fs}/LOGS",
"hf_token_file": "{shared_fs}/.cache/huggingface/token"
},
"container": {
"runtime": {
"args": {
"volumes": ["{paths.models_dir}:/models"]
}
}
}
}
Execution backends#
Four different things in this stack are called a “backend”. They are unrelated, and confusing them is the most common configuration mistake.
Setting |
Values |
What it selects |
|---|---|---|
|
|
The client backend passed to |
|
|
How vLLM distributes the model across nodes. This is the multinode setting. |
|
|
The container runtime. |
Cluster file |
|
Whether CVS runs commands on the host or inside a container. See Cluster Validation Suite (CVS) cluster file: configuration and backend selection. |
Warning
enroot is registered but not implemented. Every method is a stub that returns failure, so a run with runtime.name: "enroot" fails at test_launch_container. Podman is not supported. Use docker.
Distributed executor: mp and ray#
Multinode runs support two executor backends. mp is the default and requires no configuration key at all.
mp (default). Used whenever server_params.distributed_executor_backend is absent. The suite injects the full distributed block into each rank’s vllm serve command:
vllm serve <model> --tensor-parallel-size <tp> --port <port> \
--node-rank <rank> --master-addr <addr> --master-port <port> \
--nnodes <n> --pipeline-parallel-size <pp> \
--distributed-executor-backend mp
Every rank above 0 additionally gets --headless. This path requires pipeline parallelism (pipeline_parallel_size greater than 1).
ray (opt-in). Selected by setting server_params.distributed_executor_backend to the exact lowercase string "ray". Other values are configuration errors. Ray takes a completely different route:
Bootstrap the cluster head:
ray start --head --port=<master_port>Bootstrap each worker:
ray start --address=<cluster-head>:<dist_init_port>Launch
vllm serveon the head node only — workers run no serve processOn teardown, broadcast
ray stopafter the process kill
Under ray, none of the mp distributed flags are emitted. --pipeline-parallel-size is added only when server_params.pipeline_parallel_size is greater than 1.
Note
Ray does not enable multinode — it relaxes the pipeline-parallelism requirement. With ray, pipeline_parallel_size of 1 is legal and is the expected configuration for pure tensor-parallel multinode serving. With mp, pipeline parallelism is mandatory.
Because only the head node serves under ray, worker ranks produce no per-rank server log. That is expected.
Topology validation rules#
These rules are enforced when the configuration file loads, before anything starts:
Condition |
Rule |
|---|---|
two or more cluster hosts, backend is not ray |
|
two or more cluster hosts, backend is ray |
|
|
The distributed suite requires more than one cluster host |
two or more cluster hosts, either backend |
|
The corresponding error messages are:
multi-host distributed execution requires pipeline_parallel_size > 1 unless using ray
pipeline_parallel_size > 1 requires a multi-host distributed suite
vllm_distributed requires container.env.NCCL_SOCKET_IFNAME on multi-host clusters
Multinode prerequisites#
Beyond the validation rules, a multinode run needs:
server_params.dist_init_port— default29501; CVS derives the head address from the cluster.container.env.NCCL_IB_HCA— the comma-separated RDMA HCA names available on every node. The packaged MI3xx configurations setrdma0throughrdma7.container.env.NCCL_SOCKET_IFNAME— the Linux netdev associated with the selected RNICs.container.env.GLOO_SOCKET_IFNAMEandTP_SOCKET_IFNAME— generally the frontend/control-plane interface.container.env.NCCL_IB_GID_INDEX— the index for the intended RoCE/IB fabric. Useshow_gidsinside the container and choose an entry available on every selected HCA and node. If that command is unavailable, inspectibv_devinfo -vand/sys/class/infiniband/<hca>/ports/<port>/gid_attrs/.
Container and Docker configuration#
The container block controls image selection, lifetime, and the docker run flags.
{
"container": {
"lifetime": "per_run",
"name": "vllm_perf_inference_rocm",
"image": "rocm/vllm:latest",
"runtime": {
"name": "docker",
"args": {
"network": "host",
"ipc": "host",
"privileged": true,
"volumes": [
"/home/{user-id}:/home/{user-id}",
"{paths.models_dir}:/models"
]
}
}
}
}
Container block keys#
The following keys are accepted inside the container block.
Key |
Default |
Description |
|---|---|---|
|
none |
Container image. Required — launch fails with |
|
|
Container name. |
|
|
One of |
|
|
Container runtime. |
|
|
Docker flags; see the table below. |
|
|
Container-level environment variables. Top level, not under |
|
absent |
Path on each host to a saved image tar to |
Warning
Put env at the container top level. Placing it under runtime.args crashes container launch: the code iterates that value as a sequence of pairs, which raises ValueError: too many values to unpack for any key longer than two characters.
Runtime arguments#
All keys under runtime.args are optional. List-valued keys append to the defaults; scalar keys override them. There is no way to remove a default device or capability.
Key |
Merge |
Default |
Emitted flag |
|---|---|---|---|
|
append |
|
|
|
append |
|
|
|
append |
|
|
|
append |
|
|
|
append |
|
|
|
append |
|
|
|
override |
|
|
|
override |
|
|
|
override |
|
|
|
n/a |
none |
Triggers |
The assembled command is:
docker run -d --name <name> <args> <image> sleep infinity
The container is a long-lived sidecar; every workload command runs through docker exec inside it. InfiniBand devices are additionally passed through by per-host shell expansion at launch time, so each node mounts the devices it actually has.
Note
--gpus is deliberately never emitted — GPU access on AMD hardware comes from the /dev/kfd and /dev/dri device mounts plus the video group.
shm_size is not supported on this path. Setting runtime.args.shm_size is silently ignored; --shm-size is never emitted.
Container lifetime#
The lifetime key controls when the container is started and stopped relative to the test lifecycle.
|
Setup behavior |
Teardown behavior |
|---|---|---|
|
Verifies a container of that name is already running; never starts one. |
No-op |
|
Force-removes any stale container of the same name, then launches. |
|
|
Attaches if running on all hosts; cold-starts if absent on all hosts; refuses on partial or failed probe. |
No-op |
Tip
With persistent, always pin container.name explicitly. The default name is derived from the image, so bumping an image tag silently abandons the old container and starts a new one.
Registry authentication#
Set runtime.args.registry to log in before pulling:
{
"registry": {
"username": "myuser",
"password_file": "/path/on/each/host/to/token",
"server": "registry.example.com"
}
}
username and password_file are required; server defaults to Docker Hub. The password is read from a file path on each remote host — there is no inline password or token key, and the login is kept out of the logs. Login is skipped entirely when image_tar is set, since a tar load never pulls.
Image resolution order at launch:
If
image_taris set and the image is absent,docker loadit.Otherwise, if
registryis set, log in.Check whether the image exists on all hosts.
If not,
docker pullit, with no retry or fallback.
Note
The image-exists check matches Repository:Tag exactly, so an image referenced without a tag or by digest never matches and is pulled on every run.
Cluster file merge#
The variant’s container block is deep-merged onto the cluster file’s block: dictionaries merge key-wise, while scalars and lists are replaced. Cluster-set values survive unless the variant sets the same key.
Warning
A container block in the variant always contributes lifetime, name, and image — including empty defaults. If your variant defines container but omits image, it overwrites a cluster-file image with an empty string and the launch fails. Set image in whichever file defines the block.
Paths#
All four keys are required.
Key |
Description |
|---|---|
|
Root of the shared filesystem, typically the anchor other paths reference. |
|
Model weight cache; exported into the server as |
|
Root for run artifacts. |
|
Path to a file containing the Hugging Face token. |
If hf_token_file does not exist, the run continues with an empty token. vLLM
configs always serve a pre-staged model mounted under paths.models_dir.
Per-cell artifacts land in:
<log_dir>/vllm/out-node<rank>/isl<isl>_osl<osl>_conc<conc>/
vllm_serve_server.log
client.log
results
Model#
Key |
Default |
Description |
|---|---|---|
|
none |
Local path, for example |
Server role#
server_params controls the vllm serve process. Harness-owned fields
are model, tensor_parallel_size, pipeline_parallel_size, port,
dist_init_port, polling controls, and distributed_executor_backend.
Every other snake-case key is passed through to vllm serve.
Key |
Default |
Description |
|---|---|---|
|
none |
Local model path supplied as the positional |
|
none |
Tensor-parallel degree. |
|
|
Pipeline-parallel degree. |
|
|
OpenAI-compatible server port. |
How server_params are flattened#
JSON value |
Emitted |
Example |
|---|---|---|
Scalar |
|
|
|
|
|
|
rejected |
Omit the setting or use vLLM’s explicit negative option. |
List |
One flag followed by its values. |
|
When server_params.max_model_len is absent, CVS derives one for each
benchmark cell from its effective parameters after sweep overrides:
ceil((ISL + OSL) * (1 + random_range_ratio)) + random_prefix_len + 8
An explicit non-null server_params.max_model_len takes precedence and
emits exactly one --max-model-len flag. An explicit null value is an
intentional opt-out: the generic option serializer emits no flag, and CVS
suppresses the fallback so vLLM uses the model or image default.
Environment variables: two mechanisms#
These are separate and are frequently confused.
|
generated per-command environment |
|
|---|---|---|
Applied by |
|
A sourced shell script inside the container after HCA discovery. |
Scope |
Every command in the container, for its whole lifetime. |
Hugging Face path/token variables and legacy network fallbacks. |
Changing it |
Requires recreating the container. |
Takes effect on the next command. |
Defaults |
|
No network overrides unless a legacy top-level field is set. |
The generated per-command environment always exports:
export HF_TOKEN=<token>
export HF_HUB_CACHE=<paths.models_dir>
Packaged configurations set the HCA, socket-interface, GID, and NCCL debug
settings in container.env so every command inherits them and topology
discovery does not overwrite them. The legacy top-level ib_hca_devices and
ib_netdev fields remain fallbacks for external configurations that omit
their corresponding container environment variables. Put other static ROCm,
NCCL, and vLLM exports in container.env.
The packaged MI3xx catalog’s AITER settings are image- and model-scoped. For
the image based on vLLM commit 4bdc8a788:
DeepSeek V4 Flash, DeepSeek V4 Pro, GLM 5.1, and GLM 5.2 set
VLLM_ROCM_USE_AITER=1,VLLM_ROCM_USE_AITER_MHA=0, andGPU_ARCHS=gfx942.Kimi K2.5 sets
VLLM_ROCM_USE_AITER=1and disablesVLLM_ROCM_USE_AITER_MHA,VLLM_ROCM_USE_AITER_FP4BMM,VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS, andVLLM_ROCM_USE_AITER_MLA.Other packaged model families carry no AITER overrides.
Do not add VLLM_USE_AITER_UNIFIED_ATTENTION or
VLLM_ROCM_USE_AITER_FUSED_MOE_A16W4: this image does not register them, so
they are warning-only and ignored. Do not substitute
VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION or other attention-backend variables
without validating both registration and call sites in the exact image.
Benchmark parameters#
benchmark_params holds client defaults. Per-cell entries in sweeps
override these values. CVS owns endpoint construction, result paths, and
percentile reporting.
Key |
Default |
Description |
|---|---|---|
|
|
Client backend for |
|
|
Server base URL. |
|
|
Total prompts per cell. |
|
|
Dataset for the load generator. |
|
|
Total prompts per cell. |
|
|
1.0 is a uniform arrival process; lower is burstier. |
|
|
Random seed. |
|
|
Arrival rate; |
|
|
Length jitter around ISL/OSL; also feeds the derived max-model-len. |
|
|
Shared prefix length. |
|
|
Tokenizer mode. |
|
|
Client completion polls before giving up. |
Tip
Arbitrary snake-case keys under benchmark_params and a cell override are
translated to vllm bench serve options. They cannot override model,
endpoint, sequence lengths, concurrency, result paths, or harness-owned
percentile reporting.
Sweep#
The sweep is an explicit list of canonical cells, not a cartesian product.
sweeps defines optional overrides and runs selects the cells to execute.
{
"sweeps": {
"ISL=1000,OSL=1000,TP=8,PP=2,CONC=16": { "num_prompts": 50 },
"ISL=1000,OSL=1000,TP=8,PP=2,CONC=32": {}
},
"runs": [
"ISL=1000,OSL=1000,TP=8,PP=2,CONC=16",
"ISL=1000,OSL=1000,TP=8,PP=2,CONC=32"
]
}
Key |
Description |
|---|---|
|
Per-cell benchmark override object. |
|
Canonical key declared in |
An undeclared, malformed, duplicated, or TP/PP-inconsistent cell key is a load-time error.
Cell keys#
Each run is one cell, identified by a canonical key used to look up thresholds:
ISL=<isl>,OSL=<osl>,TP=<tp>,PP=<pp>,CONC=<conc>
PP= is always present, including single-node and Ray runs with PP=1.
The host count is placement information, not a threshold dimension.
Examples:
ISL=1000,OSL=1000,TP=8,PP=1,CONC=16
ISL=1000,OSL=1000,TP=8,PP=2,CONC=16
Server reuse#
Cells with identical server arguments share a server identity, so the suite
reuses the running server instead of stopping it, restarting, and reloading
weights. The derived --max-model-len is part of that identity: different
derived values force a restart, while cells with equal derived values reuse the
server. For example, with zero range ratio and prefix, 1024/8192 and 8192/1024
both derive 9224 and can share. An explicit
non-null server_params.max_model_len can also allow different ISL/OSL
cells to share. Explicit null removes the max-model-len option from the server
identity entirely, so different ISL/OSL cells share when their other server
arguments match. Concurrency remains client-only, so ordering runs with
concurrency varying fastest makes a sweep substantially quicker.
Thresholds#
Thresholds are keyed by canonical cell, then by a bare metric name:
{
"ISL=1000,OSL=1000,TP=8,PP=1,CONC=16": {
"output_throughput": {"kind": "min", "value": 4000},
"mean_ttft_ms": {"kind": "max", "value": 500},
"failed": {"kind": "max", "value": 0}
}
}
A cell may contain any subset of the registry, including an empty object. When
enforce_thresholds is true, every selected run must have a threshold cell,
but only specs present in that cell create verification subtests. When it is
false, produced and configured values remain record rows and no threshold
subtests run.
Sweep specs are strict at load time regardless of enforcement:
A spec contains exactly
kindandvalue.kindmust equal the registry direction, exactlyminormax.valuemust be a finite JSON number. Booleans, strings, null, arrays, objects, NaN, and infinity are rejected.Cell-level keys beginning with
_commentor_exampleare metadata and are ignored by metric validation; all other keys must name a registered metric.Prefixed names (
client.*,gpu.*,prom.*), unknown names, legacy kinds (includingmin_tok_sandmax_ms), references, tolerances, units,info, and extra fields are rejected.
min fails below its value and passes at or above it. max fails above
its value and passes at or below it. A gated actual that is missing, null,
boolean, string, collection, NaN, or infinity fails for every datasource.
The optional top-level accuracy block remains task-qualified and is exempt
from these vLLM sweep-name and spec rules.
Threshold file discovery#
Set threshold_json to the threshold file. A relative path resolves against
the configuration file’s directory. The packaged 28 files contain two cells
each and intentionally retain the previous 32-metric assertion subset in
canonical registry order. All 56 registry metrics remain supported, and every
finite produced value is reported whether or not it has a threshold spec. The
packaged zero values are uncalibrated placeholders, and every paired config
keeps enforcement disabled. Users may add any other registered metric or
remove any packaged entry.
Metrics#
vLLM uses one ordered 56-entry bare-name registry. The metric name owns its
raw source or derivation inputs, display unit, datasource, display category,
and exact threshold direction. median_* and p50_* remain separate.
Run and health (7)#
max_concurrency, max_concurrent_requests, num_prompts,
completed, failed, success_rate, duration.
Throughput and totals (10)#
request_throughput, goodput, output_throughput,
total_token_throughput, per_gpu_throughput,
decode_throughput_p50, max_output_tokens_per_s, rtfx,
total_input_tokens, total_output_tokens.
TTFT (8)#
mean_ttft_ms, median_ttft_ms, std_ttft_ms, p50_ttft_ms,
p90_ttft_ms, p95_ttft_ms, p99_ttft_ms,
normalized_ttft_ms_per_tok.
TPOT (7)#
mean_tpot_ms, median_tpot_ms, std_tpot_ms, p50_tpot_ms,
p90_tpot_ms, p95_tpot_ms, p99_tpot_ms.
ITL (8)#
mean_itl_ms, median_itl_ms, std_itl_ms, p50_itl_ms,
p90_itl_ms, p95_itl_ms, p99_itl_ms, decode_latency_ratio.
End-to-end latency (7)#
mean_e2el_ms, median_e2el_ms, std_e2el_ms, p50_e2el_ms,
p90_e2el_ms, p95_e2el_ms, p99_e2el_ms.
GPU (5)#
peak_gpu_memory_mb, model_load_memory_mb, model_load_s,
gpu_bandwidth_util_pct, gpu_compute_util_pct. Load time and memory are
captured once after successful readiness and reused for cells sharing that
server. Load time remains available whenever the elapsed measurement is finite.
Load memory requires finite pre/post VRAM snapshots; missing snapshots do not
disable server reuse or the elapsed measurement. A real zero memory delta
remains zero.
Prometheus (4)#
queue_time_p50_ms, queue_time_p95_ms, prefill_time_p50_ms, and
prefill_time_p95_ms. CVS diffs before/after histogram scrapes so reused
server counters remain isolated to one cell.
Directions#
The following are min: max_concurrency,
max_concurrent_requests, num_prompts, completed, success_rate,
request_throughput, goodput, every token throughput/total metric,
rtfx, gpu_bandwidth_util_pct, and gpu_compute_util_pct. Every other
registry metric is max.
Projection and derivation#
The vLLM result projector accepts only finite built-in integer and float values;
booleans are not numeric. Known metadata is ignored: date,
endpoint_type, backend, label, model_id, tokenizer_id,
burstiness, and request_rate (including a finite request rate). Any
unknown finite top-level numeric field fails parsing and names the artifact.
The raw request_goodput field maps only to goodput.
Derived metrics are emitted only when their result is finite:
per_gpu_throughput = total_token_throughput / (tp * pp)
normalized_ttft_ms_per_tok = mean_ttft_ms / isl
decode_latency_ratio = p99_itl_ms / p50_itl_ms
decode_throughput_p50 = 1000 / median_tpot_ms
success_rate = completed / (completed + failed)
Reporting and compatibility#
test_verify_cell_metrics remains one parent per cell. It computes all host
rows before emitting one subtest for each present, enforced spec, so one failure
does not hide sibling verdicts. Finite produced values without a spec are
record-only HTML rows. The parent also emits one compact JUnit property with
actuals_by_host and metric contract {"id":"vllm-bare","version":1}.
Run Deck tables, charts, and highlights use selected registry metrics rather than all 56 columns. Historical namespaced vLLM Run Deck artifacts do not match the contract: previous-run and manually selected viewer baselines display an explicit incompatibility and suppress performance comparisons. A compatible task-qualified accuracy section in the same baseline remains independently eligible for accuracy comparison.
Results table#
The summary table emits seven fixed columns — Model, GPU, ISL, OSL, Policy, Conc, Host — followed by Req/s, Total tok/s, Mean TTFT, P95 TTFT, Mean TPOT, P95 TPOT, P99 ITL, and Goodput.
Accuracy tests#
Accuracy evaluation runs lm-evaluation-harness against the live server after the performance sweep. Task selection lives in the configuration file; gating values live in the threshold file.
{
"accuracy": {
"tasks": [
{
"id": "gsm8k_strict",
"tasks": "gsm8k",
"backend": "vllm",
"lm_eval_model": "local-completions",
"num_fewshot": 5,
"batch_size": "auto",
"limit": 100,
"num_concurrent": 8,
"exec_timeout_sec": 7200,
"extra_model_args": "tokenizer_backend=huggingface"
}
]
}
}
Key |
Default |
Description |
|---|---|---|
|
none |
Unique label for this entry; duplicates are rejected. |
|
none |
lm-eval task name or task list. |
|
endpoint-derived |
|
|
lm-eval default |
Few-shot example count, if explicitly set. |
|
|
Concurrent requests. |
|
|
Enables or names the chat template. |
|
|
Passed through to lm-eval. |
|
|
Directory of custom task definitions. |
|
|
Generation arguments. |
lm_eval_model selects the API surface. If omitted, CVS derives it from
apply_chat_template:
Value |
lm-eval model |
Endpoint |
|---|---|---|
|
|
|
|
|
|
CVS pins runtime installation to lm-eval[api,math]==0.4.12. The shared
accuracy schema also exposes that release’s evaluation controls, including
batch_size, max_batch_size, device, limit, samples,
use_cache, cache_requests, check_integrity,
system_instruction, fewshot_as_multiturn, predict_only, seed,
trust_remote_code, confirm_run_unsafe_code, metadata, and
gen_kwargs. CVS owns the endpoint, model path, output path, and sample
logging. Results land under <log_dir>/accuracy.
Accuracy metric keys#
Accuracy metrics are keyed <lm_task_name>.<metric_key>, with any comma in the name replaced by a double underscore. For example, gsm8k’s exact_match,strict-match becomes:
gsm8k.exact_match__strict-match
Gate them in the threshold file’s top-level accuracy block, keyed by the task id:
{
"accuracy": {
"gsm8k_strict": {
"gsm8k.exact_match__strict-match": { "kind": "min", "value": 0.75 }
}
}
}
Note
An accuracy failure does not mark the shared lifecycle as failed, so the remaining stages — the results table and teardown — still run normally.
Troubleshooting#
Message |
Cause and fix |
|---|---|
Distributed execution requires |
Multi-host |
|
Use |
|
Set all three socket-interface variables under |
|
|
|
|
|
The same cell key appears twice in |
|
A |
|
A sweep or run key is malformed. |
|
The TP or PP in a cell key does not match |
|
Two |
|
The cell key does not begin with |
|
|
|
The gated datasource did not produce a valid value. Inspect the benchmark artifact, GPU telemetry, or server metrics for that cell; percentile collection is harness-owned. |
|
Remote model download is unimplemented. Pre-stage weights and set |
|
|
Extra-key validation error |
A misspelled key. Every block except |