Glossary#

Alphabetical reference for terms used in Primus documentation and configuration. Cross-links point to other production docs where applicable.


AINIC#

AMD AI NIC—AMD’s AI-optimized network interface for multi-node GPU communication (for example, the AMD Pensando™ Pollara 400 AI NIC).


Backend#

A training framework integrated into Primus (for example Megatron-LM, TorchTitan, MaxText, Megatron Bridge, HummingbirdXT).


BackendAdapter#

Abstract class in Primus that connects a backend: discovery of setup paths, config conversion, and trainer loading.


BackendRegistry#

Registry mapping backend names to adapter classes, often with lazy import to avoid loading unused frameworks.


BaseTrainer#

Abstract trainer defining the lifecycle: setupinittraincleanup.


BF16 / FP16 / FP8 / FP4#

Floating-point precisions: Brain Float 16, IEEE half, 8-bit float, and 4-bit float training or inference formats (exact support depends on backend and hardware).


CP (context parallelism)#

Parallelism that splits the sequence dimension across devices for long-context training.


DP (data parallelism)#

Replicates the model across GPUs; each rank processes different data batches.


DeepEP#

Deep Expert Parallelism—Primus-Turbo’s acceleration path for MoE token dispatch and related expert-parallel work.


EP (expert parallelism)#

Distributes Mixture-of-Experts expert networks across devices.


Experiment config#

Top-level YAML describing work_group, modules, and overrides for a training run.


FSDP#

Fully Sharded Data Parallel—shards parameters, gradients, and optimizer states across devices (PyTorch FSDP and similar concepts per backend).


GBS (global batch size)#

Total batch size across all data-parallel ranks for one optimizer step (may combine micro-batching and gradient accumulation).


Gradient accumulation#

Accumulates gradients over multiple micro-batches before an optimizer update.


HipBLASLt#

AMD’s high-performance BLAS library with autotuning for GEMM and related kernels.


Hook#

Shell or Python scripts under runner/helpers/hooks/ executed at defined lifecycle points.


LoRA#

Low-Rank Adaptation—parameter-efficient fine-tuning that trains small adapter matrices.


MBS (micro batch size)#

Batch size per GPU (per rank) for one forward/backward pass within a gradient-accumulation window.


MLA (multi-latent attention)#

Compressed KV-cache attention architecture used in models such as DeepSeek.


MoE (mixture of experts)#

Architecture with multiple expert sub-networks and a router that assigns tokens to experts.


Model config#

YAML preset describing architecture (hidden size, layers, attention heads, and so on).


Module config#

YAML preset for training behavior: learning rate, batch sizes, optimizer, schedules.


NCCL / RCCL#

NVIDIA Collective Communications Library / ROCm equivalent—libraries for GPU collective operations in distributed training.


PP (pipeline parallelism)#

Splits model layers into stages on different devices.


Patch#

Runtime monkey-patch registered in PatchRegistry and applied at a named training phase.


PatchRegistry#

Registry of phase-aware patches (for example build_args, setup, before_train, after_train).


Platform config#

YAML describing cluster environment mappings (for example platform_azure.yaml): env vars, paths, and scheduler hints.


Preflight#

Cluster diagnostic tooling that checks host, GPU, network, and baseline performance before long jobs. See primus/tools/preflight/ in the repository.


Preset#

Reusable YAML fragment under primus/configs/ (module, model, or platform).


PrimusRuntime#

Core orchestrator: loads configuration, resolves the backend, applies patches, and drives the trainer lifecycle.


Primus-SaFE#

Stability and Fault-tolerance Engine—external ecosystem component for cluster management and resilience. This repository references it in auxiliary tooling but does not include a production integration guide.


Primus-Turbo#

High-performance operator library (for example FlashAttention-style kernels, GEMM, collectives, grouped GEMM).


Projection#

Tools that estimate memory and training performance without requiring a full production cluster.


ROCm#

Radeon Open Compute—AMD’s GPU computing platform (drivers, compilers, libraries).


SFT (supervised fine-tuning)#

Supervised fine-tuning that typically updates all (or a defined subset of) model parameters, as opposed to adapter-only methods.


SP (sequence parallelism)#

Parallelism that extends tensor-parallel regions to non-TP parts of the model to reduce activation memory.


TP (tensor parallelism)#

Splits layer weights across GPUs within a node (or defined process group).


Transformer engine (TE)#

Library stack for FP8 and related training optimizations (availability depends on backend and build).


VPP (virtual pipeline parallelism)#

Interleaved pipeline parallelism with multiple virtual stages per device to improve utilization.


Zero-bubble#

Pipeline scheduling that reduces or eliminates pipeline bubbles (idle time between micro-batches).