Run ATOM LLM inference benchmarks with CVS#
2026-09-24
7 min read time
The ATOM suite benchmarks LLM serving on AMD Instinct GPUs using the ATOM stack
(atom.entrypoints.openai_server + atom.benchmarks.benchmark_serving on
single-node variants, or a PP coordinator with params.driver: vllm_atom /
sglang on multinode stems). One suite name — atom — covers single-node
and multinode pipeline-parallel runs; topology comes from the config and cluster
file.
This page walks through a first run. For the full schema, profiles, thresholds, and parameter reference, see ATOM inference benchmark configuration file for Cluster Validation Suite (CVS).
Prerequisites#
The following prerequisites are required.
On every GPU node:
Docker with ROCm device passthrough (
sudo dockeror docker group).ATOM (or vLLM-ATOM / SGLang) container image available on the node.
Model weights at
paths.models_dirwhenmodel.remote: 0.Shared log path mounted when using multinode PP.
On the launcher (where you run cvs run):
CVS installed (see Install Cluster Validation Suite (CVS) on ROCm).
SSH key access to cluster nodes (
priv_key_filein the cluster file).Hugging Face token file at
paths.hf_token_filewhen required.
For multinode PP (params.nnodes: 2, pipeline_parallel_size: 2):
Two hosts in
node_dictmatchingparams.nnodes.params.master_addrset to the head node VPC IP.IB/socket fabric discoverable, or explicit
roles.server.ib_hca_devices/roles.server.ib_netdev(socket netdev must not be anmlx5_*HCA name).
Step 1: Copy config and cluster files#
List shipped ATOM configs:
cvs config list inference/atom
Copy each variant into its own subdirectory so only one *threshold.json
sits beside the config you pass to --config_file:
SINGLE_DIR=~/input/config_file/inference/atom/single
mkdir -p "$SINGLE_DIR"
cvs config copy inference/atom/mi3xx_atom_deepseek-r1_fp8_single.json \
--output "$SINGLE_DIR/mi3xx_atom_deepseek-r1_fp8_single.json"
cvs config copy inference/atom/mi325x_atom_deepseek-r1_fp8_single_threshold.json \
--output "$SINGLE_DIR/mi325x_atom_deepseek-r1_fp8_single_threshold.json"
cvs config copy cluster_file/atom_cluster.json --output ~/input/cluster_file/atom_cluster.json
Config stems use mi3xx_* (MI300-family). Threshold files use mi325x_*
on MI325X (gfx942) — match the platform you calibrated on.
Step 2: Edit placeholders#
Replace cluster node IPs and trim node_dict to one host for single-node runs
(two hosts for distributed PP). In the config, set at minimum:
container.image— your ATOM ROCm image.container.runtime.args.volumes— replace<changeme-models-mount>with the host models directory (for example/it-share-prj2-1/modelson prj2 lab nodes).paths.shared_fs,paths.log_dir,paths.hf_token_file.model.id— model under test.
paths.models_dir is /models, the in-container mount point exported as
HF_HUB_CACHE. Keep it as shipped; only customize the host side of the models
volume mount.
For multinode PP, also set params.master_addr and verify
roles.server.ib_netdev ("auto" is the default on shipped distributed stems).
Tip
Leave enforce_thresholds: false on first lab runs until thresholds are
calibrated for your hardware. MI355X shipped stems often ship record-only.
Step 3: Run the suite#
Single-node W1 (driver=atom):
cvs run atom \
--cluster_file ~/input/cluster_file/atom_cluster.json \
--config_file "$SINGLE_DIR/mi3xx_atom_deepseek-r1_fp8_single.json" \
--html ~/cvs_results/atom-w1-single.html --self-contained-html -vvv
Multinode PP (driver=vllm_atom):
cvs run atom \
--cluster_file ~/input/cluster_file/atom_cluster.json \
--config_file ~/input/config_file/inference/atom/distributed/mi3xx_atom_deepseek-r1_fp8_distributed.json \
--html ~/cvs_results/atom-w1-distributed.html --self-contained-html -vvv
MTP-3 speculative decode (schema_version: 2 profile on the native single-node
stem):
cvs run atom \
--cluster_file ~/input/cluster_file/atom_cluster.json \
--config_file "$SINGLE_DIR/mi3xx_atom_deepseek-r1_fp8_single.json" \
--config_profile mtp3 \
--html ~/cvs_results/atom-w1-mtp3.html --self-contained-html -vvv
vLLM / SGLang parity use the unified serving schema in inference/atom/
(for example mi3xx_atom_vllm_deepseek-r1_fp8_single.json) — still run with
cvs run atom, not cvs run vllm or cvs run sglang.
Smoke one cell with pytest -k, for example -k "w1_1k_1k-conc128".
After git pull, run make install before source .cvs_venv/bin/activate.
Test lifecycle#
Tests run in a fixed order. [cell] = one row per sweep cell; [cell-tier] =
one row per metric tier per cell.
Order |
Test |
Purpose |
|---|---|---|
1 |
|
Launch and verify the container |
2 |
|
Multinode SSH setup (skipped on single-node) |
3 |
|
Resolve IB HCAs and socket netdev |
4 |
|
Verify model cache on nodes |
5 |
|
Server start (or reuse), bench client, parse results |
6 |
|
Threshold PASS/FAIL per metric tier |
7 |
|
Console summary tables |
8 |
|
Tear down the container |
On inference failure, lifecycle.failed skips downstream cells and metric
rows. When reuse_server_across_sweep: true, a warm server is kept across cells
with the same server_session_key.
Sweeps and metrics#
Each sweep cell is one (ISL, OSL, concurrency) pair from sweep.runs.
Parametrize IDs look like w1_1k_1k-conc128 or w1_1k_1k-conc128-throughput.
Threshold cell keys (must match the sibling threshold file):
Single-node:
ISL=1024,OSL=1024,TP=8,PP=1,CONC=128Multinode PP:
ISL=1024,OSL=1024,TP=8,PP=2,CONC=128
Each test_cell_metrics[cell-tier] gates one tier when enforce_thresholds:
true:
Tier |
Example metrics (bare names in thresholds) |
|---|---|
|
|
|
|
|
|
|
|
|
Remaining metrics — logged, not gated |
Benchmark artifacts still expose client.* keys internally; threshold JSON
uses bare metric names. ATOM may omit tail percentiles even when
metric_percentiles requests them — only present metrics are gated.
Reports and logs#
pytest HTML — one row per lifecycle stage and per metric tier.
Console tables —
test_print_results_tableprints per-cell throughput and latency.Run Deck — when
--htmlis set,atom_run_deck.html/.json/_viewer.htmlare bundled into the pytest zip (render-only; does not affect gates).Per-cell logs — under
paths.log_diron cluster nodes (server + client logs).
Launcher vs GPU node#
Item |
Launcher |
GPU node |
|---|---|---|
|
Yes |
No |
|
Yes |
No |
Models host mount ( |
No |
Yes |
Container image, |
No |
Yes |
|
No |
Yes |