TorchTitan training configuration files for Cluster Validation Suite (CVS)#
2026-09-24
6 min read time
TorchTitan configs live under cvs/input/config_file/training/torchtitan/. Each config has a sibling *_threshold.json referenced by threshold_json. One config file can hold multiple precision sweeps (BF16, FP8, MXFP8, MXFP4) for the same model.
Keys prefixed with _ (for example _scaling_baseline_comment) are inline comments and are ignored by the loader.
Use cvs config list training/torchtitan to list templates, or cvs config copy training/torchtitan/<name> to copy one to your working directory.
Note
Replace every
<changeme>placeholder before running; unresolved placeholders cause a hard exit at startup.{user-id}in path fields is resolved to the cluster username at load time.
Available configurations#
Templates are organized by GPU SKU; select the section that matches your hardware.
MI355X#
Config |
Model |
Mode |
|---|---|---|
|
Llama-3.1-8B |
single |
|
Llama-3.3-70B |
single |
|
Llama-3.3-70B |
distributed |
|
Llama-3.1-405B |
distributed |
|
DeepSeek-V2-Lite |
single |
|
Qwen3-32B |
single |
|
Mixtral-8x22B |
single |
MI3XX#
Config |
Model family |
Mode |
|---|---|---|
|
Llama |
single / distributed |
|
DeepSeek |
single / distributed |
|
Qwen3 |
single / distributed |
Cluster-specific edits#
Where |
Field |
Change to |
|---|---|---|
|
|
Your TorchTitan ROCm image tag, accessible on all nodes. |
|
|
Path to your Hugging Face token file on the nodes. |
|
|
Replace |
|
|
Node count and head-node IP (distributed only). |
|
|
Your NIC family and RDMA device names (distributed only). |
|
|
Measured single-node total tok/s; |
Threshold JSON |
per-metric bounds |
Calibrated PASS/FAIL limits for your hardware. |
Top-level fields#
Field |
Example |
Description |
|---|---|---|
|
|
Config schema version. Must be |
|
|
Selects the test suite. Must match |
|
|
GPU architecture label (informational). |
|
|
If |
|
|
Companion threshold file in the same directory. |
|
(block) |
Milestone-step loss decrease validation. |
|
(block) |
Steps/time to reach a target loss (informational). |
|
(block) |
Optional save/resume checkpoint test. |
|
(block) |
Single-node baseline for scaling efficiency (distributed only). |
|
(block) |
Runtime paths, NCCL, and training settings. |
|
(block) |
Model architecture defaults; sweep combos override fields. |
|
(block) |
Docker container settings. |
|
(block) |
Training combinations and ordered run list. |
config block#
Field |
Default |
Description |
|---|---|---|
|
|
Hugging Face access token file path. |
|
|
Training log output directory. |
|
|
Per-node wrapper scripts directory. |
|
|
Tokenizer and dataset cache directory. |
|
|
TorchTitan path inside the container. |
|
|
Training iterations per combo. |
|
|
Node count; must match the cluster file. |
|
|
Head-node IP for distributed coordination. |
|
|
NIC family; |
|
|
Comma-separated RDMA HCA list. |
|
|
Control-plane interface name. |
|
|
Auto-generate TorchTitan TOML from |
model_params block#
Defaults applied to every sweep combo; individual combos override matching keys.
Field |
Description |
|---|---|
|
Friendly name used in log paths and labels. |
|
Hugging Face repo ID (for example |
|
Sequence length in tokens. |
|
Dataset name (for example |
|
Learning rate and warmup steps. |
|
Parallelism degrees (TP, PP, CP, EP for MoE). |
|
Data parallel (FSDP) degree. |
|
Default precision; overridden per sweep combo. |
container block#
Field |
Description |
|---|---|
|
|
|
Docker container name. |
|
Required — Docker image URI. |
|
Host paths volume-mounted into the container. |
|
Host devices exposed ( |
Distributed configs additionally mount the Broadcom RDMA library and expose /dev/infiniband/rdma_cm.
Optional blocks#
loss_curveSample loss every
sample_everysteps and check a decreasing trend atmilestone_steps. Setenforce: trueto gate PASS/FAIL.convergenceTrack steps and wall-clock time to reach
target_valueloss. Informational only; never gates PASS/FAIL.checkpointWhen
enforce: true, run a two-phase save/resume test with step-counter and loss-continuity validation.scaling_baseline(distributed only)tokens_per_sec_totalis your measured single-node total tok/s. Used to computetraining.scaling_efficiency_pct.
Sweeps#
Each entry in sweep.combinations is one parametrized training run. sweep.runs is the ordered list of combo IDs to execute; omit combos from runs to skip them without editing combinations.
Combo field |
Description |
|---|---|
|
Human-readable label (used in reports). |
|
Batch sizes for this combo. |
|
Precision override ( |
Threshold files#
Threshold files map each sweep combo to per-metric pass/fail limits. Cell keys use the format MBS=<mbs>,GBS=<gbs>,PRECISION=<precision>.
Metric |
Description |
|---|---|
|
TFLOP/s per GPU. |
|
Tokens per second (total across all GPUs). |
|
Wall time per training step (ms). |
|
GPU memory usage. |
|
Multi-node scaling efficiency vs single-node baseline (distributed only). |
Threshold kinds: min (actual ≥ value), max (actual ≤ value), min_ratio (actual / reference ≥ value).
How to run: Run TorchTitan pre-training benchmarks with CVS.