Run SGLang LLM inference benchmarks with CVS#

2026-09-24

4 min read time

Applies to Linux

CVS provides three SGLang suites under cvs/tests/inference/sglang/. Each suite is a separate pytest module; pick the one that matches your topology, then point --config_file at a template from cvs/input/config_file/inference/sglang/.

For the full configuration schema, threshold format, and parameter reference, see SGLang inference benchmark configuration for Cluster Validation Suite (CVS).

Test suites#

CVS suite name

Source module

What it runs

sglang_single

sglang_single.py

One unified sglang.launch_server on a single benchmark_serv_node (local TP).

sglang_distributed

sglang_distributed.py

One unified multi-node server (TP/PP across server_node_list).

sglang_disagg_distributed

sglang_disagg_distributed.py

Disaggregated prefill/decode with a proxy router.

Set up config#

  1. List available SGLang templates:

    cvs config copy --list | grep inference/sglang
    
  2. Copy the configuration (and threshold file, if you edit thresholds locally):

    cvs config copy inference/sglang/mi3xx_sglang_llama_70b_single.json \
      --output ~/cvs_workspace/mi3xx_sglang_llama_70b_single.json
    
    cvs config copy inference/sglang/mi325_sglang_llama_70b_threshold.json \
      --output ~/cvs_workspace/mi325_sglang_llama_70b_threshold.json
    
  3. Copy a cluster file (container backend recommended):

    cvs config copy cluster_container.json --output ~/cvs_workspace/cluster.json
    
  4. Edit the config — set container.image, replace every <changeme> with cluster-specific values, and ensure threshold_json resolves to your threshold JSON.

Shipped config templates:

Config file

Use with suite

mi3xx_sglang_llama_70b_single.json

sglang_single

mi3xx_sglang_deepseek_r1_0528_single.json

sglang_single

mi3xx_sglang_llama_70b_distributed.json

sglang_distributed

mi3xx_sglang_deepseek_r1_0528_distributed.json

sglang_distributed

mi3xx_sglang_llama_70b_disaggregated.json

sglang_disagg_distributed

mi3xx_sglang_deepseek_r1_0528_disaggregated.json

sglang_disagg_distributed

Note

Shipped templates are a single workload per file. sweep.runs selects which ISL,OSL,TP,PP,CONC cells to parametrize; empty runs uses every performance cell in the threshold JSON.

Run tests#

List stages in a suite:

cvs list sglang_single

sglang_single stages#

Available tests in sglang_single:
  - test_launch_container
  - test_rms_norm
  - test_launch_server
  - test_poll_for_server_ready
  - test_openai_compatible_http_endpoints
  - test_run_lm_eval_hellaswag_benchmark_test
  - test_run_lm_eval_gsm8k_benchmark_test
  - test_run_performance_benchmark_test
  - test_verify_dmesg_after_benchmark
  - test_print_results_table
  - test_teardown

test_run_performance_benchmark_test is parametrized once per combo in sweep.runs (for example isl1024-osl1024-c64). Empty runs uses every performance cell in the threshold file. TP/PP in the combo may differ from the threshold-file key; matching is exact, then unique ISL,OSL,CONC.

Example run:

cvs run sglang_single \
  --cluster_file ~/cvs_workspace/cluster.json \
  --config_file ~/cvs_workspace/mi3xx_sglang_llama_70b_single.json \
  --html ~/cvs_results/sglang_single.html --self-contained-html \
  --log-file /tmp/sglang.log -vvv

sglang_distributed stages#

Available tests in sglang_distributed:
  - test_launch_container
  - test_setup_ibv_devices
  - test_rms_norm
  - test_launch_server
  - test_poll_for_server_ready
  - test_openai_compatible_http_endpoints
  - test_run_lm_eval_hellaswag_benchmark_test
  - test_run_lm_eval_gsm8k_benchmark_test
  - test_run_performance_benchmark_test
  - test_verify_dmesg_after_benchmark
  - test_distributed_gpu_topology
  - test_print_results_table
  - test_teardown

Example run:

cvs run sglang_distributed \
  --cluster_file ~/cvs_workspace/cluster.json \
  --config_file ~/cvs_workspace/mi3xx_sglang_llama_70b_distributed.json \
  --html ~/cvs_results/sglang_distributed.html --self-contained-html \
  --log-file /tmp/sglang.log -vvv

sglang_disagg_distributed stages#

Available tests in sglang_disagg_distributed:
  - test_launch_container
  - test_setup_ibv_devices
  - test_rms_norm
  - test_launch_prefill_servers
  - test_launch_decode_servers
  - test_poll_for_server_ready
  - test_launch_proxy_router
  - test_openai_compatible_http_endpoints
  - test_run_lm_eval_hellaswag_benchmark_test
  - test_run_lm_eval_gsm8k_benchmark_test
  - test_run_performance_benchmark_test
  - test_verify_dmesg_after_benchmark
  - test_disagg_gpu_topology
  - test_print_results_table
  - test_teardown

Example run:

cvs run sglang_disagg_distributed \
  --cluster_file ~/cvs_workspace/cluster.json \
  --config_file ~/cvs_workspace/mi3xx_sglang_llama_70b_disaggregated.json \
  --html ~/cvs_results/sglang_disagg.html --self-contained-html \
  --log-file /tmp/sglang.log -vvv

Direct pytest invocation#

Each module can also be run with pytest:

pytest cvs/tests/inference/sglang/sglang_single.py \
  --cluster_file ~/cvs_workspace/cluster.json \
  --config_file ~/cvs_workspace/mi3xx_sglang_llama_70b_single.json \
  --html ~/cvs_results/sglang_single.html

Read the results#

With --html, CVS writes an HTML report plus sglang_run_deck.html (interactive viewer) using the shared sglang report profile.

Key lifecycle stages to watch:

  • Container launch — test_launch_container must pass before any server work runs.

  • IB setup — test_setup_ibv_devices (distributed and disaggregated only) validates RDMA inside the container.

  • Server ready — test_poll_for_server_ready waits for the SGLang server log to show ready.

  • Smoke — test_openai_compatible_http_endpoints probes the OpenAI-compatible API.

  • Performance — test_run_performance_benchmark_test runs sglang.bench_serving for each selected sweep cell. Set top-level enforce_thresholds: false to record metrics without failing on uncalibrated gates.

  • Accuracy — test_run_lm_eval_hellaswag_benchmark_test and test_run_lm_eval_gsm8k_benchmark_test run lm-eval tasks configured in accuracy.tasks.

  • Summary — test_print_results_table prints throughput/latency/accuracy in the console and report.

  • Teardown — test_teardown stops containers even when a prior stage failed.

Logs are written under paths.log_dir from the config (default /home/{user-id}/LOGS/sglang).