Megatron training configuration files for Cluster Validation Suite (CVS)#
2026-09-24
20 min read time
JSON configs and sibling *_threshold.json files for megatron_single and megatron_distributed. MI300X and MI325X share one config per model and mode (mi3xx_megatron_{model}_{single|distributed}.json); set gpu_name and threshold_json to the SKU. MI355X ships Llama 3.1 8B and Llama 3.3 70B single-node configs only (mi355x_megatron_llama-3.1-8b_single.json, mi355x_megatron_llama-3.3-70b_single.json). Use a *_single.json file with megatron_single and a *_distributed.json file with megatron_distributed. threshold_json is resolved relative to the config file.
See Run Megatron Llama and DeepSeek training benchmarks for more information on running these tests.
Use cvs config list training/megatron to list available templates, or
cvs config copy training/megatron/<name> to copy one to your working
directory. Copy the matching SKU *_threshold.json into the same directory as
the suite config (mi300x_* or mi325x_* for mi3xx_ templates;
mi355x_* for MI355X).
Backends#
The suite selects the training backend from container.image (substring primus, case-insensitive):
Megatron-LM — image name does not contain
primus. Training scripts live under/workspace/Megatron-LMinside the image. Log files use<log_dir>/megatron-logs/<combo_id>/out-node<N>/training.log.Primus — image name contains
primus. In-image YAML lives underexamples/megatron/configs/{gpu_arch}/. Log files use<log_dir>/primus-logs/<combo_id>/out-node<N>/training.log.
Llama 3.1 8B, Llama 3.3 70B, and DeepSeek V2 Lite run on both Megatron-LM and Primus (same JSON; pick the backend with container.image). Llama 3.1 405B is Primus only (distributed).
<combo_id> on disk is the sweep combination key with non-filename characters replaced (= and , become _). Pytest still uses the unsanitized key (for example MBS=4,GBS=128,PRECISION=FP8).
Note
Parameters with the
<changeme>value must have that value modified to your specifications. Unresolved placeholders cause a hard exit at load time.{user-id}will be resolved to the cluster username (or the local OS user as fallback). You can also set this value yourself.Keys prefixed with
_(for example_checkpoint_comment) are inline comments and are ignored by the loader.sweep.runsis required whensweep.combinationsis non-empty. It must be a non-empty subset of (or equal to) those keys; an emptyrunslist is a load error. Omittingsweepentirely (or using emptycombinations) runs one implicitdefaultcell;train_params.micro_batch_size,global_batch_size, andprecisionare then required.Each key in
sweep.combinationsmust beMBS=<micro_batch_size>,GBS=<global_batch_size>,PRECISION=<precision>. The suite parses those three values from the key; do not repeat them in the combination body. The body may overlay extras such astraining_iterations,tensor_parallelism, andpipeline_parallelism.
Available configurations#
MI300X and MI325X share mi3xx_megatron_<model>_<mode>.json. MI355X ships mi355x_megatron_llama-3.1-8b_single.json and mi355x_megatron_llama-3.3-70b_single.json. Thresholds stay SKU-specific (mi300x_*_threshold.json, mi325x_*_threshold.json, mi355x_*_threshold.json) and are selected with threshold_json.
The table below maps each supported model to its configuration file and available execution modes.
Model |
MI300X / MI325X config |
MI355X config |
Mode |
|---|---|---|---|
Llama 3.1 8B |
|
|
single, distributed (Megatron-LM or Primus); MI355X packaged as single-node |
Llama 3.3 70B |
|
|
single, distributed (Megatron-LM or Primus); MI355X packaged as single-node |
DeepSeek V2 Lite |
|
— |
single, distributed (Megatron-LM or Primus) |
Llama 3.1 405B |
|
— |
distributed (Primus only) |
Single-node configs set container.env.MASTER_ADDR to 127.0.0.1. NNODES is not in the JSON: the suite sets it from the cluster host count into docker run -e. Distributed configs require MASTER_ADDR and NCCL_IB_HCA and add a scaling_baseline section and checkpoint_dir to the checkpoint block. NIC type is not in the JSON: Megatron-LM defaults to thor2 for MI300X/MI325X and ainic for MI355X from gpu_name.
Leftover files mi3xx_megatron_llama_*.json and mi35x_megatron_llama_single.json are nested-schema configs for the legacy megatron_llama3_1_* suites only. Do not pass them to megatron_single or megatron_distributed.
Required edits#
Set these before a run (full field tables are under Common parameters):
The following fields must be customized to match your cluster and hardware before invoking any Megatron suite.
gpu_name/threshold_json— onmi3xx_templates, setMI300XorMI325Xand the matching SKU threshold filename.gpu_namemust be exactlyMI300X,MI325X, orMI355Xafter uppercase.container.image— Megatron-LM or Primus ROCm image on all nodes.train_params.training_iterations— default training steps (for example"30"). Eachsweep.combinationsbody setstraining_iterations(packaged value"20"); that overlay wins for that cell.train_params.micro_batch_size/global_batch_size/precision— required whensweepis omitted. Packaged files already set them; with a declared sweep the combination key supplies those values for each cell.paths.hf_token_file— Hugging Face token path on the nodes.container.env.NCCL_SOCKET_IFNAME/GLOO_SOCKET_IFNAME/NCCL_IB_GID_INDEX/NCCL_DEBUG— templates include example values plus<changeme>(for exampleenp193s0f1np1 <changeme>,3 <changeme>,ERROR <changeme>). Remove<changeme>and keep or edit the example. Required on single-node and distributed.sweep.runs— combination keys to execute (sameMBS=…,GBS=…,PRECISION=…strings as insweep.combinations). Must be non-empty whencombinationsis present.Do not set
NNODESin JSON.Distributed only:
container.env.MASTER_ADDR,container.env.NCCL_IB_HCA(example HCA list plus<changeme>).paths.data_cache_dirmust be a shared filesystem. Whencheckpoint.enforceistrue, also setcheckpoint.checkpoint_dirand replace the last<changeme>:<changeme>volume with that shared path.
Top-level fields#
These fields appear at the root of every config file.
Field |
Example |
Description |
|---|---|---|
|
|
GPU architecture string used for logging and Primus YAML path resolution ( |
|
Host paths: |
|
|
Image, runtime mounts, and |
|
|
see per-model tables |
Model knobs plus |
|
omitted (single) / |
Compare RDMA and ethtool error counters before and after training. Single-node templates omit it (schema/lib default |
|
|
If |
|
|
Filename of the companion threshold file, looked up in the same directory as the config file. On |
Model configurations#
Each model section below shows only the train_params and sweep blocks, which are the parts that differ between models. All other sections (paths, container, smoke, loss_curve, convergence, checkpoint, scaling_baseline) are identical in structure across models and are documented in Common parameters.
Llama 3.1 8B#
Available as mi3xx_megatron_llama-3.1-8b_{single,distributed}.json (MI300X/MI325X) and mi355x_megatron_llama-3.1-8b_single.json (MI355X).
mi3xx_megatron_llama-3.1-8b_single.json (representative)
{
"gpu_name": "<changeme>",
"enforce_thresholds": true,
"threshold_json": "<changeme>",
"train_params": {
"model_name": "llama3.1_8B",
"tokenizer_model": "meta-llama/Llama-3.1-8B",
"model_size": "8",
"sequence_length": "8192",
"recompute": "0",
"fsdp": "0",
"tensor_parallelism": "1",
"pipeline_parallelism": "1",
"micro_batch_size": "4",
"global_batch_size": "128",
"precision": "BF16",
"training_iterations": "<changeme>"
},
"sweep": {
"combinations": {
"MBS=4,GBS=128,PRECISION=FP8": {
"training_iterations": "20"
},
"MBS=4,GBS=128,PRECISION=BF16": {
"training_iterations": "20"
},
"MBS=4,GBS=128,PRECISION=MXFP4": {
"training_iterations": "20"
},
"MBS=4,GBS=128,PRECISION=MXFP8": {
"training_iterations": "20"
}
},
"runs": [
"MBS=4,GBS=128,PRECISION=FP8",
"MBS=4,GBS=128,PRECISION=BF16",
"MBS=4,GBS=128,PRECISION=MXFP4",
"MBS=4,GBS=128,PRECISION=MXFP8"
]
}
}
train_params#
Parameter |
Value |
Description |
|---|---|---|
|
|
Used in log labels and report filenames. |
|
|
HuggingFace repo ID for the tokenizer and model weights. |
|
|
Model size in billions of parameters. |
|
|
Maximum context length. |
|
|
Fits on a single GPU; no tensor splitting needed. |
|
|
Single pipeline stage. |
|
|
Required when |
|
|
Required when |
|
|
Required when |
|
|
Training steps (same field on every model). |
Llama 3.3 70B#
Available as mi3xx_megatron_llama-3.3-70b_{single,distributed}.json (MI300X/MI325X) and mi355x_megatron_llama-3.3-70b_single.json (MI355X).
mi3xx_megatron_llama-3.3-70b_single.json (representative)
{
"gpu_name": "<changeme>",
"enforce_thresholds": true,
"threshold_json": "<changeme>",
"train_params": {
"model_name": "llama3.3_70B",
"tokenizer_model": "meta-llama/Llama-3.3-70B-Instruct",
"model_size": "70",
"sequence_length": "8192",
"recompute": "0",
"fsdp": "0",
"tensor_parallelism": "8",
"pipeline_parallelism": "1",
"micro_batch_size": "3",
"global_batch_size": "96",
"precision": "BF16",
"training_iterations": "<changeme>"
},
"sweep": {
"combinations": {
"MBS=3,GBS=96,PRECISION=FP8": {
"training_iterations": "20"
},
"MBS=3,GBS=96,PRECISION=BF16": {
"training_iterations": "20"
}
},
"runs": [
"MBS=3,GBS=96,PRECISION=FP8",
"MBS=3,GBS=96,PRECISION=BF16"
]
}
}
train_params#
Parameter |
Value |
Description |
|---|---|---|
|
|
Used in log labels and report filenames. |
|
|
HuggingFace repo ID for the tokenizer and model weights. |
|
|
Model size in billions of parameters. |
|
|
Maximum context length. |
|
|
Splits the model across all 8 GPUs on the node. |
|
|
Single pipeline stage. |
|
|
Required when |
|
|
Required when |
|
|
Required when |
|
|
Training steps (same field on every model). |
DeepSeek V2 Lite#
Available as mi3xx_megatron_deepseek-v2-lite_{single,distributed}.json (MI300X/MI325X).
mi3xx_megatron_deepseek-v2-lite_single.json (representative)
{
"gpu_name": "<changeme>",
"enforce_thresholds": true,
"threshold_json": "<changeme>",
"train_params": {
"model_name": "deepseek_v2_lite",
"tokenizer_model": "deepseek-ai/DeepSeek-V2-Lite",
"model_size": "16",
"sequence_length": "4096",
"recompute": "0",
"fsdp": "0",
"tensor_parallelism": "1",
"pipeline_parallelism": "1",
"micro_batch_size": "4",
"global_batch_size": "128",
"precision": "BF16",
"training_iterations": "<changeme>"
},
"sweep": {
"combinations": {
"MBS=4,GBS=128,PRECISION=BF16": {
"training_iterations": "20"
},
"MBS=4,GBS=128,PRECISION=FP8": {
"training_iterations": "20"
}
},
"runs": [
"MBS=4,GBS=128,PRECISION=BF16",
"MBS=4,GBS=128,PRECISION=FP8"
]
}
}
train_params#
Parameter |
Value |
Description |
|---|---|---|
|
|
Used in log labels and report filenames. |
|
|
HuggingFace repo ID. CVS downloads the |
|
|
Model size in billions of parameters. |
|
|
Shorter context than Llama due to DeepSeek V2’s attention architecture. |
|
|
Fits on a single GPU. |
|
|
Single pipeline stage. |
|
|
Required when |
|
|
Required when |
|
|
Required when |
|
|
Training steps (same field on every model). |
Llama 3.1 405B#
Available as mi3xx_megatron_llama-3.1-405b_distributed.json (MI300X/MI325X; distributed only). Unlike the other models, 405B runs on Primus only: container.image must contain primus. Megatron-LM does not support 405B.
mi3xx_megatron_llama-3.1-405b_distributed.json (representative)
{
"gpu_name": "<changeme>",
"enforce_thresholds": true,
"threshold_json": "<changeme>",
"train_params": {
"model_name": "llama3.1_405B",
"tokenizer_model": "meta-llama/Llama-3.1-405B",
"model_size": "405",
"sequence_length": "8192",
"recompute": "0",
"fsdp": "0",
"tensor_parallelism": "8",
"pipeline_parallelism": "4",
"micro_batch_size": "1",
"global_batch_size": "64",
"precision": "BF16",
"training_iterations": "<changeme>"
},
"sweep": {
"combinations": {
"MBS=1,GBS=64,PRECISION=FP8": {
"training_iterations": "20"
},
"MBS=1,GBS=64,PRECISION=BF16": {
"training_iterations": "20"
}
},
"runs": [
"MBS=1,GBS=64,PRECISION=FP8",
"MBS=1,GBS=64,PRECISION=BF16"
]
}
}
train_params#
Parameter |
Value |
Description |
|---|---|---|
|
|
Used in log labels and report filenames. |
|
|
HuggingFace repo ID for the tokenizer and model weights. |
|
|
Model size in billions of parameters. |
|
|
Maximum context length. |
|
|
Splits across all 8 GPUs per node. |
|
|
Splits the model across 4 pipeline stages (requires at least 4 nodes). |
|
|
Required when |
|
|
Required when |
|
|
Required when |
|
|
Training steps. |
Common parameters#
These sections appear in all config files. The parameter names and semantics are identical across models and GPU variants.
paths#
Host paths used by the job. They must be volume-mounted into the container (typically via the home-directory bind mount).
Parameter |
Default |
Description |
|---|---|---|
|
|
Path to a Hugging Face token file for gated models and datasets. |
|
|
Host path where per-node training logs are written. Megatron-LM writes |
|
|
Host path where the lib writes per-rank wrapper scripts. |
|
|
Dataset and tokenizer cache directory. On distributed runs this path must be a shared filesystem. |
|
|
ROCm installation path inside the container. Leave empty for auto-detection. |
container.env#
These values are passed into the container as docker run -e flags and also flattened into the training job dict. NNODES is injected at launch from the cluster host count (not a JSON field). Packaged templates put example NIC strings in the values and keep <changeme> in the same string so load fails until you edit them. Megatron-LM and Primus wrappers do not re-export NCCL_*, GLOO_SOCKET_IFNAME, MASTER_ADDR, or NNODES. Megatron-LM still re-exports NCCL_IB_GID_INDEX after Broadcom NIC setup.
Parameter |
Default |
Description |
|---|---|---|
|
|
Rank-0 address (loopback on single-node). |
|
|
Network interface for NCCL control channels. Replace |
|
|
Network interface for Gloo control channels. Same example pattern as |
|
|
GID index for InfiniBand addressing. Replace |
|
|
NCCL log verbosity. Replace |
|
|
(Distributed) Comma-separated InfiniBand HCA device names. The packaged list is an example; replace |
container#
Parameter |
Default |
Description |
|---|---|---|
|
|
When to create and destroy the container. |
|
(model-specific) |
Container instance name, e.g. |
|
|
Docker image to run. If the image name contains |
|
|
Container runtime. Currently only |
|
|
Use host networking so NCCL and Gloo can reach other nodes directly. |
|
|
Share the host IPC namespace for GPU shared memory. |
|
|
Required for ROCm GPU and InfiniBand device access. |
|
(see below) |
List of |
|
|
GPU device nodes to expose. Distributed configs also add |
checkpoint#
Controls the checkpoint save and resume test (test_checkpoint). The test is Primus-only: it is skipped when enforce is false, and also skipped when the container image name does not contain primus. On Primus it runs in two phases: a save phase that trains for save_iters steps writing a checkpoint every save_interval steps, followed by a resume phase that loads the last checkpoint and trains to resume_iters steps. Continuity is checked at the first resume step (last_ckpt_step + 1), which must not exceed the checkpoint-step loss by more than loss_rtol. Micro-batch size, global batch size, and precision are the same as test_smoke (the smoke block).
checkpoint_dir is only present in distributed configs. On Primus distributed runs it must be a shared filesystem path visible on every node. Single-node Primus ignores that field and writes under {log_dir}/ckpt_primus. Megatron-LM never runs test_checkpoint.
Load I/O timing is taken from the node-0 Primus log. Single-node resume lines say loading checkpoint from; distributed resume lines say loading distributed checkpoint from. Both end with successfully loaded checkpoint from. A missing load-start line yields a warning, not a test failure.
Parameter |
Default |
Description |
|---|---|---|
|
|
If |
|
|
How often (in steps) to write a checkpoint during the save phase. The last checkpoint lands at |
|
|
Steps to train in the save phase. Must not be an exact multiple of |
|
|
Steps to train in the resume phase, continuing from the last checkpoint. |
|
|
Relative tolerance for the loss continuity check. The first step of the resume phase must not exceed the checkpoint-step loss by more than |
|
|
(Distributed only) Shared filesystem path for checkpoints. Must be volume-mounted into the container at the same path on all nodes. Required only when |
smoke#
Controls test_smoke: a small fixed cell (not a sweep.runs entry) that loads the model and trains a few steps with no metric gating. test_checkpoint uses the same micro_batch_size, global_batch_size, and precision (it still uses checkpoint.* for step counts). Packaged configs set this explicitly; if the block is omitted the loader defaults to enabled with iters 10, MBS 1, precision BF16, and an empty global_batch_size (the suite then uses 8 on single-node and 16 on distributed).
Parameter |
Default |
Description |
|---|---|---|
|
|
If |
|
|
Training steps for the smoke cell. |
|
|
Micro-batch size for the smoke cell and for |
|
|
Global batch size for smoke and |
|
|
Precision tag for the smoke cell and for |
loss_curve#
Parameter |
Default |
Description |
|---|---|---|
|
|
Sample a loss point every N steps for the slope check. |
|
|
Additional steps always included in the sampled loss curve regardless of |
|
|
Maximum allowed least-squares slope of the sampled loss curve. A positive slope (loss increasing) fails the check. |
|
|
If |
convergence#
Parameter |
Default |
Description |
|---|---|---|
|
|
Metric tracked for convergence. |
|
|
Loss value at which the model is considered converged. |
scaling_baseline#
(Distributed configs only.)
Parameter |
Default |
Description |
|---|---|---|
|
|
Total tokens/sec from a prior single-node run ( |
|
|
Number of nodes used to produce |
sweep#
Parameter |
Default |
Description |
|---|---|---|
|
N/A |
Dict of sweep cells. The key must be |
|
|
Per-cell step count. Overrides |
|
N/A |
Ordered list of combination keys to execute. Required and non-empty when |
Pytest parametrizes sweep_name from sweep.runs. The suite parses micro_batch_size, global_batch_size, and precision from each combination key. Additional fields in the combination body override the corresponding train_params values for that run, including training_iterations, tensor_parallelism, and pipeline_parallelism. If sweep is omitted or combinations is empty, one cell named default trains with train_params; micro_batch_size, global_batch_size, and precision must be set there. The matching threshold cell is default and is required when enforce_thresholds is true.
Threshold files#
Each suite JSON names a sibling file in threshold_json. With a declared sweep, top-level cell keys must be the same MBS=<mbs>,GBS=<gbs>,PRECISION=<precision> strings as sweep.runs (unused combinations keys are not required). With no sweep (generic / implicit default run), the threshold cell is the literal key default — not an MBS=… string built from train_params. A missing or extra cell fails load when enforce_thresholds is true. When it is false, the default cell may be omitted: the loader warns and test_metric is record-only.
A metric is gated only when enforce_thresholds is true and the cell has a numeric spec:
Kind |
Passes when |
|---|---|
|
actual ≥ value |
|
actual ≤ value |
|
always; recorded only |
|
actual / |
|
if true, Megatron |
Tracked metrics (namespace training.*): throughput_per_gpu, tokens_per_gpu, elapsed_time_per_iteration, mem_usage (Megatron-LM mem usages:; Primus does not emit it; packaged specs set optional: true), and on distributed configs scaling_efficiency_pct (kind: info and optional: true in packaged files).