SGLang inference benchmark configuration for Cluster Validation Suite (CVS)#
2026-09-24
9 min read time
CVS ships three SGLang inference suites for AMD MI30X clusters. Each suite reads a JSON
configuration file from cvs/input/config_file/inference/sglang/ and a matching
threshold file referenced by top-level threshold_json.
CVS suite |
Test module |
Topology |
|---|---|---|
|
|
One unified |
|
|
One unified multi-node server (TP/PP + |
|
|
Disaggregated prefill/decode with a proxy router; separate prefill and decode node groups. |
See Run SGLang LLM inference benchmarks with CVS for more information on running these tests.
Run any suite with:
cvs run <suite> \
--cluster_file cvs/input/cluster_file/<cluster>.json \
--config_file cvs/input/config_file/inference/sglang/<config>.json \
--html=~/cvs_results/sglang.html
Copy a template locally:
cvs config list inference/sglang
cvs config copy inference/sglang/mi3xx_sglang_llama_70b_single.json \
--output ~/cvs_workspace/inference/sglang/mi3xx_sglang_llama_70b_single.json
cvs config copy inference/sglang/mi325_sglang_llama_70b_threshold.json \
--output ~/cvs_workspace/inference/sglang/mi325_sglang_llama_70b_threshold.json
Note
{user-id}in path strings is resolved to the current username at runtime.Replace every
<changeme>placeholder before running; unresolved placeholders cause a hard exit at startup.
Configuration files#
All SGLang templates live under cvs/input/config_file/inference/sglang/.
Model and topology templates#
Config file |
Threshold JSON |
Use with suite |
|---|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Threshold files#
Performance cells and pass/fail limits are stored separately. Each workload points at its
threshold file via top-level threshold_json (a filename beside the config).
Threshold file |
Referenced by |
|---|---|
|
Llama 3.1 70B configs (single, distributed, disaggregated) |
|
DeepSeek-R1-0528 configs (single, distributed, disaggregated) |
Threshold keys use the form ISL=<n>,OSL=<n>,TP=<n>,PP=<n>,CONC=<n>. Each value is a
metric map (for example output_throughput_per_sec, mean_ttft_ms, mean_tpot_ms,
goodput, mfu) with kind and value fields. Accuracy cells use
BENCH=lm_eval_hellaswag and BENCH=lm_eval_gsm8k.
File structure#
Shipped templates use these top-level keys:
The following table describes each top-level key in a SGLang configuration file.
Key |
Description |
|---|---|
|
When |
|
Filename of the threshold JSON in the same directory. |
|
|
|
Image, name, lifetime, and |
|
Model, TP/PP, node lists, ports, and launch flags. |
|
|
|
|
|
Optional per-combo overrides (for example |
|
|
The HuggingFace or filesystem path loaded by SGLang is server_params.model.
Example: single-node template#
mi3xx_sglang_llama_70b_single.json (abbreviated)
{
"enforce_thresholds": false,
"threshold_json": "mi325_sglang_llama_70b_threshold.json",
"paths": {
"shared_fs": "/home/{user-id}",
"models_dir": "/root/models",
"log_dir": "{shared_fs}/LOGS/sglang",
"hf_token_file": "{shared_fs}/.hf_token"
},
"container": {
"lifetime": "per_run",
"name": "sglang_container",
"image": "<changeme>",
"runtime": {
"name": "docker",
"args": {
"network": "host",
"ipc": "host",
"privileged": true,
"shm_size": "128G",
"volumes": [
"/home/{user-id}:/home/{user-id}",
"/mnt/dtni/models:/root/models"
],
"devices": [ "/dev/dri", "/dev/kfd" ],
"env": {
"NCCL_DEBUG": "ERROR",
"SGLANG_USE_AITER": "1"
}
}
}
},
"server_params": {
"backend": "sglang",
"nnodes": "1",
"model": "meta-llama/Llama-3.1-70B-Instruct",
"tensor_parallelism": "8",
"pipeline_parallelism": "1",
"benchmark_serv_node": "<changeme>",
"proxy_router_serv_port": "8000",
"add_flags": [ "--attention-backend aiter" ]
},
"benchmark_params": {
"backend": "sglang",
"data_set_name": "random",
"num_prompts": "25",
"model_num_params": "70000000000",
"peak_gpu_tflops": "2615"
},
"accuracy": { "tasks": [ { "id": "lm_eval_hellaswag" }, { "id": "lm_eval_gsm8k" } ] },
"sweep": {
"runs": [
{ "combo": "ISL=1024,OSL=1024,TP=8,PP=1,CONC=64" }
]
}
}
General config parameters#
Parameter |
Example |
Description |
|---|---|---|
|
|
Docker image with SGLang and ROCm for MI30X. |
|
|
Container instance name on each participating node. |
|
|
Server rank count. For |
|
|
HuggingFace token file for model download. |
|
|
Docker shared memory size. |
|
|
Shared log root (must be visible from benchmark nodes). |
|
|
SGLang server log level. |
|
|
NCCL log level (multi-node). |
|
node hostname/IP |
Node that runs smoke tests, lm-eval, and |
|
|
HTTP port for the unified server (single/distributed) or proxy router client port (disaggregated). |
|
|
GPU devices passed into the container. Multi-node configs also include |
|
list of |
Bind mounts for home, models, and (multi-node) RDMA libraries. See Volume mounts. |
Single-node only (sglang_single)#
Parameter |
Description |
|---|---|
|
Exactly one host; only this node receives a container. Other cluster nodes are ignored. |
|
Must be |
Unified multi-node (sglang_distributed)#
Additional server_params / container env fields beyond the single-node set:
Parameter |
Description |
|---|---|
|
All ranks of the unified |
|
Distributed init port on rank-0 (default |
|
NCCL InfiniBand/RoCE device list and GID index ( |
|
Ethernet interfaces for socket/Gloo fallback. |
|
Single |
Disaggregated prefill-decode (sglang_disagg_distributed)#
Uses the multi-node network env fields above, plus:
Parameter |
Description |
|---|---|
|
Node groups for prefill and decode servers. |
|
Host running the PD proxy router. |
|
Internal service ports (defaults |
|
Rank-0 addresses for each role group. |
|
Coordinator ports (defaults |
benchmark_params / model settings#
Inference tests#
benchmark_paramsRandom synthetic load via
sglang.bench_serving. ISL/OSL/concurrency cells come fromsweep.runs(or every performance cell in the threshold file ifrunsis empty).input_lengthandoutput_lengthare injected at collection time.Combo keys are matched exactly, then by unique
ISL,OSL,CONCif TP/PP in the combo differs from the threshold-file key.sweepssupplies per-combo overrides such asnum_prompts.enforce_thresholds(top-level): whenfalse, measured throughput/latency is recorded and reported but does not fail the run. Whentrue, results are compared against the matched threshold cell.num_prompts,random_range_ratio,model_num_params,peak_gpu_tflops: bench workload and MFU calculation inputs.
accuracy.tasks(lm_eval_hellaswag,lm_eval_gsm8k)Accuracy tasks via lm-eval. Thresholds for accuracy metrics are always enforced when configured in the threshold file.
Volume mounts#
Single-node configs mount only user home and model storage—no InfiniBand or RDMA verb libraries.
Distributed and disaggregated configs add RDMA-related mounts for Thor/Broadcom NICs:
{
"volumes": [
"/dev/infiniband:/dev/infiniband",
"/usr/local/lib/libbnxt_re-rdmav34.so:/usr/lib/x86_64-linux-gnu/libibverbs/libbnxt_re-rdmav34.so:ro",
"/usr/lib/x86_64-linux-gnu/libibverbs.so.1:/usr/lib/x86_64-linux-gnu/libibverbs.so.1:ro",
"/lib/libibverbs.d:/lib/libibverbs.d"
]
}
test_setup_ibv_devices (distributed and disaggregated suites only) validates IB visibility inside
the container after these mounts are applied.
Disaggregated architecture overview#
SGLang disaggregated prefill-decode separates inference into:
Prefill nodes — process prompts and build KV cache.
Decode nodes — autoregressive token generation from cached KV states.
Proxy router — routes requests between prefill and decode clusters.
Use sglang_disagg_distributed with mi3xx_sglang_*_disaggregated.json templates. Unified
multi-node serving (no PD split) uses sglang_distributed instead.
Performance metrics#
The results table and threshold files use:
Output throughput (
output_throughput_per_sec) — output tokens per second.TTFT (
mean_ttft_ms) — mean time to first token.TPOT (
mean_tpot_ms) — mean time per output token.E2E latency (
mean_e2e_latency_ms) — end-to-end request latency.Goodput — fraction of successful requests.
MFU — model FLOPs utilization derived from
model_num_paramsandpeak_gpu_tflops.
Troubleshooting#
- Container launch
Verify
container.imageon all nodes,devicesGPU paths, andshm_size. Single-node runs need only/dev/driand/dev/kfd. Put ROCm/SGLang knobs as scalarenvkeys (SGLANG_USE_AITER,GPU_ARCHS, …), not as anADD_EXPORT_ENVlist.- Multi-node networking
Confirm RDMA devices with
ibv_devinfoinside the container aftertest_setup_ibv_devices. MatchNCCL_IB_HCAto your cluster. For Thor NICs, ensurelibbnxt_re-rdmav34.somounts are present.- Sweep collection
Each listed
sweep.runscombo must match a threshold cell exactly or uniquely byISL,OSL,CONC. Emptyrunsselects every performance cell in the threshold JSON.- Model access
Set
paths.hf_token_filefor HuggingFace models or mount local weights under/root/modelsviavolumes.