Run SGLang LLM inference benchmarks with CVS#
2026-09-24
4 min read time
CVS provides three SGLang suites under cvs/tests/inference/sglang/. Each suite is a
separate pytest module; pick the one that matches your topology, then point --config_file
at a template from cvs/input/config_file/inference/sglang/.
For the full configuration schema, threshold format, and parameter reference, see SGLang inference benchmark configuration for Cluster Validation Suite (CVS).
Test suites#
CVS suite name |
Source module |
What it runs |
|---|---|---|
|
|
One unified |
|
|
One unified multi-node server (TP/PP across |
|
|
Disaggregated prefill/decode with a proxy router. |
Set up config#
List available SGLang templates:
cvs config copy --list | grep inference/sglang
Copy the configuration (and threshold file, if you edit thresholds locally):
cvs config copy inference/sglang/mi3xx_sglang_llama_70b_single.json \ --output ~/cvs_workspace/mi3xx_sglang_llama_70b_single.json cvs config copy inference/sglang/mi325_sglang_llama_70b_threshold.json \ --output ~/cvs_workspace/mi325_sglang_llama_70b_threshold.json
Copy a cluster file (container backend recommended):
cvs config copy cluster_container.json --output ~/cvs_workspace/cluster.json
Edit the config — set
container.image, replace every<changeme>with cluster-specific values, and ensurethreshold_jsonresolves to your threshold JSON.
Shipped config templates:
Config file |
Use with suite |
|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
Note
Shipped templates are a single workload per file. sweep.runs selects which
ISL,OSL,TP,PP,CONC cells to parametrize; empty runs uses every performance
cell in the threshold JSON.
Run tests#
List stages in a suite:
cvs list sglang_single
sglang_single stages#
Available tests in sglang_single:
- test_launch_container
- test_rms_norm
- test_launch_server
- test_poll_for_server_ready
- test_openai_compatible_http_endpoints
- test_run_lm_eval_hellaswag_benchmark_test
- test_run_lm_eval_gsm8k_benchmark_test
- test_run_performance_benchmark_test
- test_verify_dmesg_after_benchmark
- test_print_results_table
- test_teardown
test_run_performance_benchmark_test is parametrized once per combo in sweep.runs
(for example isl1024-osl1024-c64). Empty runs uses every performance cell in the
threshold file. TP/PP in the combo may differ from the threshold-file key; matching is
exact, then unique ISL,OSL,CONC.
Example run:
cvs run sglang_single \
--cluster_file ~/cvs_workspace/cluster.json \
--config_file ~/cvs_workspace/mi3xx_sglang_llama_70b_single.json \
--html ~/cvs_results/sglang_single.html --self-contained-html \
--log-file /tmp/sglang.log -vvv
sglang_distributed stages#
Available tests in sglang_distributed:
- test_launch_container
- test_setup_ibv_devices
- test_rms_norm
- test_launch_server
- test_poll_for_server_ready
- test_openai_compatible_http_endpoints
- test_run_lm_eval_hellaswag_benchmark_test
- test_run_lm_eval_gsm8k_benchmark_test
- test_run_performance_benchmark_test
- test_verify_dmesg_after_benchmark
- test_distributed_gpu_topology
- test_print_results_table
- test_teardown
Example run:
cvs run sglang_distributed \
--cluster_file ~/cvs_workspace/cluster.json \
--config_file ~/cvs_workspace/mi3xx_sglang_llama_70b_distributed.json \
--html ~/cvs_results/sglang_distributed.html --self-contained-html \
--log-file /tmp/sglang.log -vvv
sglang_disagg_distributed stages#
Available tests in sglang_disagg_distributed:
- test_launch_container
- test_setup_ibv_devices
- test_rms_norm
- test_launch_prefill_servers
- test_launch_decode_servers
- test_poll_for_server_ready
- test_launch_proxy_router
- test_openai_compatible_http_endpoints
- test_run_lm_eval_hellaswag_benchmark_test
- test_run_lm_eval_gsm8k_benchmark_test
- test_run_performance_benchmark_test
- test_verify_dmesg_after_benchmark
- test_disagg_gpu_topology
- test_print_results_table
- test_teardown
Example run:
cvs run sglang_disagg_distributed \
--cluster_file ~/cvs_workspace/cluster.json \
--config_file ~/cvs_workspace/mi3xx_sglang_llama_70b_disaggregated.json \
--html ~/cvs_results/sglang_disagg.html --self-contained-html \
--log-file /tmp/sglang.log -vvv
Direct pytest invocation#
Each module can also be run with pytest:
pytest cvs/tests/inference/sglang/sglang_single.py \
--cluster_file ~/cvs_workspace/cluster.json \
--config_file ~/cvs_workspace/mi3xx_sglang_llama_70b_single.json \
--html ~/cvs_results/sglang_single.html
Read the results#
With --html, CVS writes an HTML report plus sglang_run_deck.html (interactive viewer)
using the shared sglang report profile.
Key lifecycle stages to watch:
Container launch —
test_launch_containermust pass before any server work runs.IB setup —
test_setup_ibv_devices(distributed and disaggregated only) validates RDMA inside the container.Server ready —
test_poll_for_server_readywaits for the SGLang server log to show ready.Smoke —
test_openai_compatible_http_endpointsprobes the OpenAI-compatible API.Performance —
test_run_performance_benchmark_testrunssglang.bench_servingfor each selected sweep cell. Set top-levelenforce_thresholds: falseto record metrics without failing on uncalibrated gates.Accuracy —
test_run_lm_eval_hellaswag_benchmark_testandtest_run_lm_eval_gsm8k_benchmark_testrun lm-eval tasks configured inaccuracy.tasks.Summary —
test_print_results_tableprints throughput/latency/accuracy in the console and report.Teardown —
test_teardownstops containers even when a prior stage failed.
Logs are written under paths.log_dir from the config (default /home/{user-id}/LOGS/sglang).