Model support matrix#
This document summarizes which model families Primus targets per backend and lists representative checked-in model presets and example experiment YAML under the repository. It distinguishes curated examples from theoretical support (a preset or upstream stack may exist without a matching examples/ entry). Use the filesystem under primus/configs/models/ and examples/*/configs/ as the authoritative live inventory.
For how to add presets, see Adding model configurations. Backend parameter references: Megatron, TorchTitan, MaxText, Megatron Bridge.
Overview: Supported model families (high level)#
The following aligns with the backend overview and the configs present in this tree.
Backend |
Model families (documentation / stack scope) |
|---|---|
Megatron-LM |
LLaMA2 / LLaMA3 / LLaMA3.1 / LLaMA3.3 / LLaMA4 (sizes from small to 405B+), DeepSeek-V2 (including lite) and DeepSeek-V3, Mixtral MoE and large MoE recipe YAML, Qwen2.5 and Qwen3 (dense and MoE), Grok, GPT-OSS (20B / 120B), GLM, Kimi K2, LFM2, MiniMax, Zebra LLaMA, Mamba, and generic |
TorchTitan |
LLaMA3 family (including 3.1), LLaMA4 examples, DeepSeek-V3 examples, and Qwen3 examples including 0.6B, 1.7B, 4B, 8B, 14B, and 32B variants where present. Additional presets exist under |
MaxText (JAX) |
LLaMA2 / LLaMA3 / LLaMA3.3, DeepSeek-V2 16B, Mixtral-8x7B, Grok1, Qwen3 14B / 30B-A3B (per presets and examples). Broader coverage may exist in upstream MaxText; see MaxText. |
Megatron Bridge |
Qwen3 pretraining and post-training examples, plus post-training examples for Zebra LLaMA and Mamba where present. LLaMA 3.1 70B Bridge examples appear under MI355X. |
HummingbirdXT |
Registered backend with a post-training trainer and one checked-in example; user-facing support level still needs maintainer confirmation. |
Interpretation: “Supported” in upstream code can exceed what this repository ships as YAML. Rows below reference representative files that exist under primus/configs/models/ and examples/; they should not be treated as a complete generated inventory.
Megatron model configs#
Model presets live in primus/configs/models/megatron/. Example experiments that reference those presets appear under examples/megatron/configs/MI300X/, MI325X/, and MI355X/.
For TorchTitan, the MI300X, MI325X, and MI355X example directories carry the same model set (21 configs each). For Megatron, MI300X and MI325X are nearly identical except that MI325X omits qwen3_5_35B_A3B (BF16 and FP8)—so MI300X has 70 example configs while MI325X has 68—and MI355X is a superset (99 configs; it adds models such as glm5, gpt_oss_120B, kimi_k2, lfm2_8B_A1B, and minimax_m2.5). Each row’s SKU list below reflects exactly which SKUs ship a curated example (see, for example, qwen3_5_35B_A3B, which is MI300X/MI355X only).
Model name (file) |
Preset path |
Role |
Example experiment dirs |
Precision in examples |
|---|---|---|---|---|
|
|
Dense model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Base fragment ( |
— |
— |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
MoE model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Base fragment |
— |
— |
|
|
Model preset |
No curated example in this repo |
— |
|
|
Model preset |
MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Base fragment |
— |
— |
|
|
Base fragment |
— |
— |
|
|
MoE model preset |
MI355X |
BF16, FP8 |
|
|
Generic Megatron LM defaults |
Used via |
— |
|
|
MoE model preset |
MI355X |
BF16, FP8 |
|
|
Base fragment |
— |
— |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Base fragment |
— |
— |
|
|
Base fragment |
— |
— |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Base fragment |
— |
— |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
MoE model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
MoE model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Base fragment |
— |
— |
|
|
Model preset |
MI300X, MI325X, MI355X |
Set in experiment overrides |
|
|
Base fragment |
— |
— |
|
|
MoE model preset |
MI355X |
BF16, FP8 |
|
|
MoE model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
MoE model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Base fragment |
— |
— |
|
|
Large MoE template |
No curated example in this repo |
— |
|
|
Large MoE template |
No curated example in this repo |
— |
|
|
Large MoE template |
No curated example in this repo |
— |
|
|
Large MoE template |
No curated example in this repo |
— |
|
|
MoE proxy / test template |
No curated example in this repo |
— |
|
|
Primus Megatron root defaults |
Used via |
— |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Base fragment |
— |
— |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
MoE model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
MoE model preset |
MI300X, MI355X |
BF16, FP8 |
|
|
MoE model preset |
MI300X, MI325X, MI355X |
BF16, FP8 |
|
|
Model preset |
MI300X, MI325X, MI355X |
Set in experiment overrides |
|
|
Model preset |
MI300X, MI325X, MI355X |
Set in experiment overrides |
|
|
Model preset |
MI300X, MI325X, MI355X |
Set in experiment overrides |
Parallelism: Tensor, pipeline, and expert parallel sizes are not fixed in model presets; they are set in experiment overrides (for example tensor_model_parallel_size, pipeline_model_parallel_size, expert_model_parallel_size). MoE presets such as qwen3_235B_A22B.yaml typically require non-default expert parallelism in real runs—see the matching experiment YAML.
TorchTitan model configs#
Presets: primus/configs/models/torchtitan/. Examples: examples/torchtitan/configs/MI300X/, MI325X/, and MI355X/.
Model name (file) |
Preset path |
Example experiment dirs |
Precision in examples |
|---|---|---|---|
|
|
MI300X, MI325X, MI355X |
BF16 |
|
|
MI300X, MI325X, MI355X |
FP8 |
|
|
MI300X, MI325X, MI355X |
BF16 |
|
|
MI300X, MI325X, MI355X |
FP8 |
|
|
MI300X, MI325X, MI355X |
(see experiment) |
|
|
Preset only; stock examples use |
— |
|
|
No example in this repo |
— |
|
|
No example in this repo |
— |
|
|
No example in this repo |
— |
|
|
No example in this repo |
— |
|
|
MI300X, MI325X, MI355X |
BF16 |
|
|
MI300X, MI325X, MI355X |
FP8 |
|
|
MI300X, MI325X, MI355X |
BF16 |
|
|
MI300X, MI325X, MI355X |
FP8 |
|
|
MI300X, MI325X, MI355X |
BF16 |
|
|
MI300X, MI325X, MI355X |
FP8 |
|
|
No example in this repo |
— |
|
|
No example in this repo |
— |
|
|
No example in this repo |
— |
|
|
MoE; MI300X, MI325X, MI355X |
BF16 |
|
|
MoE; MI300X, MI325X, MI355X |
FP8 |
|
|
MoE; MI300X, MI325X, MI355X |
BF16 |
|
|
MoE; MI300X, MI325X, MI355X |
FP8 |
|
|
MI300X, MI325X, MI355X |
(see experiment) |
|
|
MI300X, MI325X, MI355X |
(see experiment) |
|
|
MI300X, MI325X, MI355X |
(see experiment) |
|
|
MI300X, MI325X, MI355X |
(see experiment) |
|
|
MI300X, MI325X, MI355X |
(see experiment) |
|
|
MI300X, MI325X, MI355X |
(see experiment) |
Parallelism: Controlled by TorchTitan launch configuration and Primus module overrides (see TorchTitan patch notes and TorchTitan parameters); not embedded in the small job / model preset alone.
MaxText model configs#
Presets: primus/configs/models/maxtext/. Examples: examples/maxtext/configs/MI300X/ and examples/maxtext/configs/MI355X/.
Model name (file) |
Preset path |
Example experiment dirs |
|---|---|---|
|
|
MI300X, MI355X |
|
|
MI300X |
|
|
MI300X, MI355X |
|
|
MI300X, MI355X |
|
|
MI300X, MI355X |
|
|
MI300X, MI355X |
|
|
MI355X |
|
|
MI300X, MI355X |
|
|
MI300X, MI355X |
|
|
MI300X, MI355X |
|
|
MI300X, MI355X |
|
|
Extended by other presets (not a standalone run) |
Parallelism: JAX / MaxText sharding is configured in experiment overrides (for example ici_fsdp_parallelism, ici_data_parallelism, dcn_* in sample experiments). See MaxText parameters.
Megatron Bridge model configs#
Presets: primus/configs/models/megatron_bridge/. Examples: examples/megatron_bridge/configs/MI300X/ and examples/megatron_bridge/configs/MI355X/.
Model name (file) |
Preset path |
Recipe / flavor (from preset) |
Example experiment dirs |
|---|---|---|---|
|
|
|
MI300X pretrain, MI355X posttrain |
|
|
|
MI300X, MI355X |
|
|
|
MI355X |
|
|
Zebra LLaMA presets |
MI300X posttrain |
|
|
Mamba preset |
MI300X posttrain |
Example filenames include *_pretrain.yaml, *_sft_posttrain.yaml, and *_lora_posttrain.yaml; precision such as bf16_mixed is set in experiment overrides.
Hardware compatibility (example directories)#
Curated example layouts under examples/ use GPU SKU subdirectories. As of this document:
GPU SKU |
|
|
|
|
|---|---|---|---|---|
MI300X |
Yes |
Yes |
Yes |
Yes |
MI355X |
Yes |
Yes |
Yes |
Yes |
MI325X |
Yes |
Yes |
No |
No |
Megatron and TorchTitan ship MI325X example directories in addition to MI300X and MI355X examples. MaxText includes MI300X and MI355X examples, including MI355X-only entries such as llama3.1_405B-pretrain.yaml. Megatron Bridge MI300X examples include Qwen3 8B and 32B pretraining plus Qwen3 32B, Zebra LLaMA, and Mamba post-training examples; LLaMA 3.1 70B Bridge examples appear under MI355X.
Absence of a SKU directory for a given backend does not imply the backend cannot run there; it means this tree does not currently provide a checked-in example path to copy from.
Model architecture reference (Megatron presets)#
Values below come from primus/configs/models/megatron/ presets (merged through extends). Vocabulary size is usually defined by the tokenizer / Hugging Face config, not duplicated in every YAML; context is max_position_embeddings where set in the chain. Use this table as a quick reference for common sizes—not an exhaustive spec of every parameter.
Model family |
Example preset |
Hidden size |
Layers |
Attention heads |
KV heads (GQA) |
Max position (context) |
|---|---|---|---|---|---|---|
LLaMA 2 7B |
|
4096 |
32 |
32 |
32 (no GQA) |
From |
LLaMA 3 8B |
|
4096 |
32 |
32 |
8 |
8192 ( |
LLaMA 3 70B |
|
8192 |
80 |
64 |
8 |
8192 |
LLaMA 3.1 405B |
|
16384 |
126 |
128 |
8 |
8192 |
Qwen3 8B |
|
4096 |
36 |
32 |
8 |
131072 ( |
Mixtral 8x7B |
|
4096 |
32 |
32 |
— |
4096 |
DeepSeek-V3 (MoE) |
|
7168 |
61 |
128 (MLA) |
— |
See preset / HF |
Mamba 370M |
|
(Mamba stack) |
— |
— |
— |
— |
For MoE and hybrid architectures (LLaMA 4, Qwen3-MoE, large moe_*.yaml templates), refer to the full YAML and upstream model cards; headline dimensions alone do not capture expert layout or MLA.