Run Megatron Llama and DeepSeek training benchmarks#
2026-09-24
15 min read time
Cluster validation that runs Megatron-LM or Primus pre-training on AMD Instinct GPUs (single-node or multi-node) and gates the run on performance and correctness metrics with a PASS/FAIL HTML report.
The suite drives a training job inside a Docker container on one or more cluster nodes, then parses the training log to produce metrics and verdicts. It provides:
Two suites —
megatron_single(single-node) andmegatron_distributed(multi-node, adds RDMA/NIC setup).Megatron-LM or Primus — if
container.imagecontainsprimus(case-insensitive), the suite uses Primus; otherwise Megatron-LM. Llama 3.1 8B, Llama 3.3 70B, and DeepSeek V2 Lite support both backends. Llama 3.1 405B is Primus only (distributed). Primus reads YAML fromexamples/megatron/configs/{gpu_arch}/inside the image.Parameter sweeps — one full training run per enabled combo (for example FP8 and BF16), each with its own result rows in the report.
Loss curve — a per-combo decreasing-trend check on
lm_lossat steps 100 / 500 / 1k / 5k.Training-log error scanning — NCCL, GPU HW faults, OOM, and other signatures fail a run early with a clear reason.
HTML report — per-test rows with linked logs and a consolidated metric results page.
The mode is which suite you invoke (megatron_single vs megatron_distributed). Use a *_single.json config with the single-node suite and a *_distributed.json config with the distributed suite.
Prerequisites#
The following prerequisites are required.
Passwordless SSH from the control host to each cluster node (key in the cluster file) and Docker available on the nodes.
A container image for ROCm (
container.imagein the config). A Megatron-LM image must provide Megatron-LM at/workspace/Megatron-LM. A Primus image (name containsprimus) uses in-image YAML underexamples/megatron/configs/{gpu_arch}/instead.A Hugging Face token file at
paths.hf_token_file(used to fetch the tokenizer). Tokenizer download requires network access on the nodes. For gated models (LLaMA, DeepSeek), model access must be granted on huggingface.co.For distributed runs: RDMA interfaces configured and reachable on all nodes; a shared filesystem path reachable from all nodes for
paths.data_cache_dir, logs, and scripts.
Set up config#
Follow these steps to set up the training configuration.
List available training configuration files:
cvs config list training/megatron
Copy the configuration file and the SKU-specific threshold file, for example:
cvs config copy training/megatron/mi3xx_megatron_llama-3.1-8b_single.json --output ~/cvs_workspace/training/megatron/mi3xx_megatron_llama-3.1-8b_single.json cvs config copy training/megatron/mi300x_megatron_llama-3.1-8b_single_threshold.json --output ~/cvs_workspace/training/megatron/mi300x_megatron_llama-3.1-8b_single_threshold.json # or mi325x_megatron_llama-3.1-8b_single_threshold.json for MI325X
Replace every
<changeme>with cluster-specific values. For MI300X/MI325X shared templates, setgpu_nametoMI300XorMI325Xandthreshold_jsonto the matchingmi300x_*ormi325x_*threshold file. Also setcontainer.imageand thecontainer.envNIC fields (templates ship example interface/HCA/GID/debug strings that still contain<changeme>— keep or edit the example and remove the placeholder). Packagedsweep.combinationscells already settraining_iterationsto"20"(that overlay wins overtrain_params.training_iterations). Change the overlay per combo if needed. Do not addNNODES; the suite sets it from the cluster host count atdocker run. Distributed configs also needMASTER_ADDR,NCCL_IB_HCA, and a sharedpaths.data_cache_dir.Change any other parameters relevant to your testing requirements.
The same folder also has DeepSeek V2 Lite (single and distributed; Megatron-LM or Primus) and Llama 3.1 405B (distributed, Primus only). See Config and threshold files for the full inventory.
Full parameter list: Megatron training configuration files for Cluster Validation Suite (CVS).
Primus vs Megatron-LM#
The same eight test stages run for both backends. _make_training_job in megatron_single.py / megatron_distributed.py picks the job class from container.image (substring primus, case-insensitive). Llama 3.1 8B, Llama 3.3 70B, and DeepSeek V2 Lite support both Megatron-LM and Primus. Llama 3.1 405B is Primus only (distributed).
Stage |
Megatron-LM (image does not contain |
Primus (image contains |
|---|---|---|
|
Launch the Docker container on all suite hosts |
Same |
|
Downloads |
Always a no-op ( |
|
Runs the cell from the config |
Same cell via |
|
Skipped ( |
Runs only when |
|
Wrapper script + Megatron-LM shell. Distributed also runs |
Wrapper script + |
|
Parse Megatron-LM log metrics |
Parse Primus log metrics (same |
|
Tear down the container |
Same |
Logs for both backends use per-node files: <log_dir>/{megatron-logs|primus-logs}/<combo_id>/out-node<N>/training.log. On disk <combo_id> is the sweep key with = and , replaced by _ (for example MBS_4_GBS_128_PRECISION_FP8).
Run tests#
You can list all available test stages using the CLI:
cvs list megatron_single
Available tests in megatron_single:
- test_launch_container
- test_download_tokenizer
- test_smoke
- test_checkpoint
- test_training
- test_metric
- test_loss_curve
- test_teardown
cvs list megatron_distributed
Available tests in megatron_distributed:
- test_launch_container
- test_download_tokenizer
- test_smoke
- test_checkpoint
- test_training
- test_metric
- test_loss_curve
- test_teardown
Use a *_single.json config with megatron_single and a *_distributed.json config with megatron_distributed.
--cluster_file— JSON describing the node(s); see Configure the Cluster Validation Suite (CVS) cluster file (cluster.json).--config_file— one of the files underinput/config_file/training/megatron/; field reference: Megatron training configuration files for Cluster Validation Suite (CVS).--html/--self-contained-html— write the HTML report.
Single-node — MI300X / MI325X#
Use the shared mi3xx_ template. Set gpu_name and threshold_json to the SKU before running.
cvs run megatron_single \
--cluster_file input/cluster_file/cluster.json \
--config_file input/config_file/training/megatron/mi3xx_megatron_llama-3.1-8b_single.json \
--html ./logs/megatron_single.html --self-contained-html -vvv -s
Single-node — MI355X#
cvs run megatron_single \
--cluster_file input/cluster_file/cluster.json \
--config_file input/config_file/training/megatron/mi355x_megatron_llama-3.1-8b_single.json \
--html ./logs/megatron_single.html --self-contained-html -vvv -s
Distributed — MI300X / MI325X#
Use the shared mi3xx_ template. Set gpu_name and threshold_json to the SKU before running.
cvs run megatron_distributed \
--cluster_file input/cluster_file/cluster.json \
--config_file input/config_file/training/megatron/mi3xx_megatron_llama-3.3-70b_distributed.json \
--html ./logs/megatron_distributed.html --self-contained-html -vvv -s
Run a specific stage#
cvs run megatron_single test_smoke \
--cluster_file input/cluster_file/cluster.json \
--config_file input/config_file/training/megatron/mi3xx_megatron_llama-3.1-8b_single.json
Test lifecycle#
Tests run in this pinned order. [combo] = one row per enabled sweep combo.
Order |
Test |
Runs on |
Purpose |
|---|---|---|---|
0 |
|
once |
Launch and verify the container |
1 |
|
once |
Megatron-LM: download |
2 |
|
once |
Small run from the |
3 |
|
once |
Primus-only checkpoint save + resume with loss continuity at the first resume step. Uses the |
4 |
|
per combo |
Build cmd, train, poll logs, parse results; GPU memory freed between combos |
5 |
|
per combo |
Threshold PASS/FAIL per metric |
6 |
|
per combo |
Gate on downward |
7 |
|
once |
Tear the container down |
A training failure is isolated to that combo’s test_training row; other combos still run. When a combo’s training does not complete, its downstream test_metric and test_loss_curve rows are skipped. If an early lifecycle stage fails, all subsequent stages are skipped via lifecycle.failed.
On a training failure, lingering GPU processes are killed (stop_training_processes) so the next combo does not launch on top of them.
Sweeps#
A sweep combo is one full training run declared in sweep.combinations. Each combination key must be MBS=<micro_batch_size>,GBS=<global_batch_size>,PRECISION=<precision>; the suite parses those values from the key, so they are not repeated in the combination body. Packaged templates use this body (other train_params overlays such as tensor_parallelism and pipeline_parallelism are optional):
"MBS=4,GBS=128,PRECISION=BF16": {
"training_iterations": "20"
}
sweep.runs is the ordered list of combination keys to execute; set it to a subset to run only selected combos without editing combinations. An empty runs list with a non-empty combinations dict fails at load.
Omitting sweep (or leaving combinations empty) runs one implicit cell named default. train_params must then set micro_batch_size, global_batch_size, and precision (load fails if any are missing). The threshold file must then have a top-level default cell when enforce_thresholds is true. When enforce_thresholds is false, that cell is optional (load warns; metrics are record-only). Packaged configs already declare a sweep and keep MBS=… threshold keys; they do not use default.
Pytest parametrizes sweep_name (one row per sweep.runs entry, or default when there is no sweep) for test_training, test_metric, and test_loss_curve. The combo ID (for example MBS=4,GBS=128,PRECISION=FP8) is that pytest ID and the threshold cell key.
Metrics and PASS/FAIL#
Each test_metric[combo] compares the parsed metric against its threshold spec and reports one of:
Status |
Meaning |
|---|---|
PASS |
value satisfies the threshold |
FAIL |
value violates the threshold (row is red; aggregated in the summary) |
RECORD |
no threshold defined, or |
SKIPPED |
spec has |
Metrics surfaced (namespace training.*):
Metric |
Description |
|---|---|
|
TFLOP/s per GPU |
|
Tokens per GPU per second |
|
Wall time per training step (ms) |
|
GPU memory usage (Megatron-LM |
|
Multi-node scaling efficiency % vs single-node baseline (distributed only). Packaged specs use |
Gating requires enforce_thresholds: true in the config. Set to false for record-only runs.
Scaling efficiency (distributed only)#
test_training computes scaling efficiency as:
efficiency % = (actual_total_tok/s / (actual_nodes / baseline_nodes)) / baseline_total_tok/s × 100
Populate scaling_baseline.tokens_per_sec_total in the config from a completed single-node run (tok/s/GPU × 8). Set to 0.0 to disable and collect data only.
Loss curve#
test_loss_curve[combo] fits a least-squares line to lm_loss samples collected at loss_curve.milestone_steps (plus every loss_curve.sample_every steps) and passes when the slope is below loss_curve.max_slope. A value of 0.0 means any downward trend passes.
Set loss_curve.enforce: false to record the slope without gating.
Convergence#
test_metric[combo] also reports a convergence check when convergence.target_value > 0. It compares the final training loss (or eval loss when eval runs) against convergence.target_value. Set target_value <= 0 to disable (record-only). target_metric: "auto" selects eval loss when available, otherwise training loss.
Checkpoint save and resume#
test_checkpoint runs only for Primus images (container.image contains primus) and only when checkpoint.enforce is true. Megatron-LM images skip this stage. Batch size and precision come from the smoke block (same resolution as test_smoke). Step counts come from checkpoint.save_iters / save_interval / resume_iters. When it runs, the suite launches the training job twice:
Save phase — trains to
checkpoint.save_iterssteps, saving a checkpoint everycheckpoint.save_intervalsteps. Single-node uses{log_dir}/ckpt_primus. Distributed usescheckpoint.checkpoint_dir(required; must be a shared path such as NFS).Resume phase — resumes from the saved checkpoint and trains to
checkpoint.resume_iterssteps.
The suite passes when:
Step counter — the resume phase starts at
last_ckpt_step + 1.Loss continuity — the loss at that first resume step does not exceed the checkpoint-step loss by more than
checkpoint.loss_rtol(relative tolerance).
Checkpoint I/O duration is parsed from the node-0 Primus training log (informational; a parse miss logs a warning and does not fail the test):
Save — timestamps on
saving checkpoint at iteration N/successfully saved checkpoint from iteration N.Load start (single-node) —
loading checkpoint from ...Load start (distributed) —
loading distributed checkpoint from ...Load end (both) —
successfully loaded checkpoint from ...
Non-zero ranks (for example rank 15 of 16) typically omit those load lines, so distributed test_checkpoint uses out-node0/training.log.
Training-log error detection#
During polling, each node’s training.log is scanned for known error patterns. Defaults cover:
NCCL errors and timeouts
GPU hardware faults and hangs
PyTorch distributed errors
A match fails that combo’s test_training with the matched pattern name and the last lines of the log.
Reports and logs#
Results table — one row per test; metric rows show PASS/FAIL from the threshold check.
Full log — each test row links to its own captured log.
Training logs — Megatron-LM writes
<log_dir>/megatron-logs/<combo_id>/out-node<N>/training.log. Primus writes<log_dir>/primus-logs/<combo_id>/out-node<N>/training.log. Folder names use the sanitized combo ID (see the table below).
Log path fields:
Placeholder |
Source |
|---|---|
|
|
|
Sweep run ID (for example |
|
One directory per node; |
Config and threshold files#
Located in cvs/input/config_file/training/megatron/. Field-level schema: Megatron training configuration files for Cluster Validation Suite (CVS).
container.env NIC fields include example values plus <changeme>. NNODES is not a JSON field. Every config except Llama 3.1 405B supports Megatron-LM or Primus; 405B is Primus only.
Do not use leftover mi3xx_megatron_llama_*.json / mi35x_megatron_llama_single.json files with these suites. Those nested configs belong only to the legacy megatron_llama3_1_* test modules.
MI355X#
Config |
Threshold |
Mode |
|---|---|---|
|
|
single-node |
|
|
single-node |
Legacy suite names#
Earlier releases shipped per-model suite names such as megatron_llama3_1_8b_single,
megatron_llama3_1_70b_distributed, and megatron_llama3_1_8b_distributed.
These suites are superseded by the unified megatron_single and megatron_distributed
suites documented above. Use the unified suites for all new deployments.