ATOM inference benchmark configuration file for Cluster Validation Suite (CVS)#
2026-09-24
8 min read time
The ATOM suite validates LLM serving on AMD Instinct GPUs. Single-node variants
use params.driver: atom (native openai_server). Shipped multinode
pipeline parallel stems use params.driver: vllm_atom or sglang with
nnodes: 2 and pipeline_parallel_size: 2.
Run the suite with:
cvs run atom --cluster_file <cluster.json> --config_file <config.json>
For a step-by-step first run, see Run ATOM LLM inference benchmarks with CVS. Config
JSON uses full param names (for example tensor_parallelism); threshold cell
keys use abbreviations (for example TP=, PP=).
File layout#
Shipped files live under cvs/input/config_file/inference/atom/:
{gpu}_atom_{model}_{precision}_single.json
{gpu}_atom_{model}_{precision}_distributed.json # when multinode PP is supported
{platform}_atom_{model}_{precision}_threshold.json
Config stems use the family prefix mi3xx. Threshold files use a
platform prefix (shipped: mi325x, lab-validated on MI325X / gfx942).
Each config’s threshold_json points at the matching mi325x_*_threshold.json.
Multi-profile configs (schema_version: 2) embed job shapes under profiles.
Select one at runtime with --config_profile NAME (or CVS_CONFIG_PROFILE).
Flat schema_version: 1 files use an implicit perf profile.
cvs.lib.utils.config_loader.substitute_config() resolves threshold_json
next to the config. Multiple *threshold.json files in one directory raises an
ambiguous-threshold error — on lab machines, copy each variant pair into its own
subdirectory (see Run ATOM LLM inference benchmarks with CVS).
Shipped model inventory#
Lab-validated configs only. Native ATOM driver=atom workloads use flat
schema_version: 1 JSON or schema_version: 2 profiles (perf, mtp3).
Framework parity (vLLM / SGLang) uses the unified serving schema under
inference/atom/ (stems mi3xx_atom_vllm_*, mi3xx_atom_sglang_*) and
still runs with cvs run atom.
Model stem |
Config files |
Notes |
|---|---|---|
|
|
Native ATOM + multinode PP |
|
|
Native ATOM perf + accuracy |
|
|
vLLM parity (serving schema) |
|
|
GPT-OSS MXFP4 vLLM parity (serving schema) |
|
|
SGLang parity (serving schema) |
Config profiles#
schema_version: 2 atom files embed job shapes under profiles. Select one at
runtime with --config_profile (or CVS_CONFIG_PROFILE). Flat
schema_version: 1 files use an implicit perf profile.
DeepSeek R1 FP8 — mi3xx_atom_deepseek-r1_fp8_single.json profiles:
|
Driver |
Use when |
|---|---|---|
|
|
Daily perf gate + full accuracy matrix |
|
|
Speculative-decode perf + gsm8k quality |
atom_vllm / atom_sglang parity (serving schema, same suite):
cvs run atom \
--config_file ~/input/.../mi3xx_atom_sglang_deepseek-r1_fp8_single.json \
--cluster_file ~/input/cluster_file/atom_cluster.json
cvs run atom \
--config_file ~/input/.../mi3xx_atom_vllm_deepseek-r1_fp8_single.json \
--cluster_file ~/input/cluster_file/atom_cluster.json
cvs run atom \
--config_file ~/input/.../mi3xx_atom_vllm_gpt-oss-120b_mxfp4_single.json \
--cluster_file ~/input/cluster_file/atom_cluster.json
cvs run atom \
--config_file ~/input/.../mi3xx_atom_deepseek-r1_fp8_single.json \
--config_profile mtp3 \
--cluster_file ~/input/cluster_file/atom_cluster.json
Keys prefixed with _ (for example _comment) are ignored by the loader.
Cluster and lab setup#
Cluster file: cvs/input/cluster_file/atom_cluster.json. Edit
node_dict so host count matches params.nnodes.
Variant type |
|
|
|---|---|---|
Single-node / baseline / MTP3 |
|
Head node only |
Multinode PP |
|
Head + worker |
Set enforce_thresholds to true for PASS/FAIL gating or false for
record-only (MI355X stems often ship record-only until lab calibration).
Fields you must customize#
Replace the following fields with values specific to your lab and cluster before running the suite.
Where |
Change to |
|---|---|
|
Your ATOM ROCm image on GPU nodes |
|
Host directory holding HF cache / weights; mounted read-only at |
|
Lab paths; |
|
Model under test |
|
|
|
Match topology ( |
|
Single-node output tok/s for |
|
Inline ATOM server CLI when |
|
Multinode server flags when |
|
|
Threshold JSON values |
Calibrated PASS/FAIL bounds for your hardware |
Placeholder substitution#
{user-id}— cluster username (or local OS user fallback).{shared_fs}— self-reference withinpaths.{paths.models_dir}and other{paths.*}— cross-referenced anywhere.<changeme-models-mount>— host path to the HF cache / weights tree; mounted at/modelsin the container. Shipped configs setpaths.models_dirto/models(the in-container path used forHF_HUB_CACHE).{head-node-ip}— replace manually in copied multinode configs.
threshold_json is a literal filename (no placeholder substitution).
Configuration schema#
Top-level fields:
Field |
Meaning |
|---|---|
|
|
|
|
|
Inferred from the config filename prefix ( |
|
|
|
Sibling threshold filename |
|
See below |
params block#
Field |
Meaning |
|---|---|
|
|
|
TP size (W1: |
|
PP size ( |
|
Node count (not part of threshold cell key) |
|
Multinode PP coordinator |
|
Benchmark client workload |
|
Tail percentiles requested (e.g. |
|
Keep server warm across cells with matching session key |
|
Baseline for |
|
Server ready and client completion timeouts |
sweep block#
Field |
Meaning |
|---|---|
|
Named |
|
|
Each cell’s threshold key is built by cvs.lib.inference.atom.atom_config_loader.AtomVariantConfig.cell_key():
Single-node:
ISL=1024,OSL=1024,TP=8,PP=1,CONC=128Multinode PP:
ISL=1024,OSL=1024,TP=8,PP=2,CONC=128
Always include PP= (use PP=1 on single-node). params.nnodes is still
required for multinode runs but is not part of the threshold key.
Execution drivers#
Driver |
When |
Server |
Native PP |
|---|---|---|---|
|
Single-node W1, baseline, MTP3 |
|
No |
|
Shipped 2-node PP stems |
Multinode serve + ATOM ROCm env |
Yes |
Multinode fabric is probed in test_discover_topology on the cluster host
OS (not inside the container).
Optional blocks#
Block |
Purpose |
|---|---|
|
FUNC-1 / FUNC-2 before sweep |
|
INF-6 dmesg / INF-7 GPU metrics |
|
lm-eval tasks (thresholds under top-level |
|
NIAH long-context cells (ACC-12) |
|
MTP acceptance checks (ACC-4/5/13) |
Threshold files#
A threshold file maps each cell key to {metric: spec}. Perf metrics use
bare names (output_throughput, mean_ttft_ms, …). Multinode cells
may also gate scaling.efficiency_pct. Non-sweep keys include accuracy,
mtp_quality, and long_context_accuracy.
Threshold kinds:
|
Passes when |
Notes |
|---|---|---|
|
|
e.g. |
|
|
Latency upper bound |
|
|
Throughput lower bound |
|
Tolerance or ratio vs reference |
Needs extra fields |
|
Always record |
Calibrate later |
Example single-node cell:
"ISL=1024,OSL=1024,TP=8,PP=1,CONC=128": {
"output_throughput": {"kind": "min_tok_s", "value": 1500},
"p99_ttft_ms": {"kind": "max_ms", "value": 5000},
"success_rate": {"kind": "min", "value": 1},
"failed": {"kind": "max", "value": 0}
}
Example multinode cell:
"ISL=1024,OSL=1024,TP=8,PP=2,CONC=128": {
"output_throughput": {"kind": "min_tok_s", "value": 2500},
"scaling.efficiency_pct": {"kind": "min", "value": 80}
}
When enforce_thresholds: true, every member of
cvs.lib.inference.atom.atom_parsing.GATED_METRICS needs a spec in each
gated cell. Metrics missing from the benchmark artifact are skipped at gate time.
Run deck comparison (optional)#
Set these env vars or sibling files for render-only comparison panels (no effect on pytest gates):
Env / file |
Panel |
Purpose |
|---|---|---|
|
|
Per-cell throughput delta vs baseline |
Baseline |
|
gsm8k flexible-extract delta |
|
|
Framework parity ratios |