AMD LLM Extension 26.09 release notes#
3 min read time
This is the seventh release of the AMD LLM Extension toolkit, an open-source software toolkit built on the ROCm platform for large language model (LLM) extensions, integrations, and performance enablement on AMD GPUs. The domain brings together training, post-training, inference, and orchestration components to make modern LLM stacks practical and reproducible on AMD hardware.
Release highlights#
Note
AMD LLM Extension 26.09 introduces updates to two components (verl and Ray) as part of the toolkit; other components remain unchanged (FlashInfer, ROCm-RAG, and Triton Inference Server).
This release updates the following components with support for ROCm 10.0.0:
verl (Volcano Engine Reinforcement Learning for LLMs) is an open-source framework for reinforcement-learning post-training of large language models and is the reference implementation of HybridFlow.
Ray is an open-source framework for scaling Python and AI workloads, providing the distributed compute and orchestration layer used by verl for multi-GPU and multi-node training.
System requirements#
AMD LLM Extension components span a range of ROCm version requirements depending on the specific extension. Ensure you follow the installation instructions for each component, which list the exact ROCm dependencies, or refer to the Compatibility matrix to verify the supported ROCm versions.
AMD LLM Extension components#
The following table lists AMD LLM Extension component versions for the 26.09 release. Click to go to the component’s source on GitHub.
| Name | Version | Source |
|---|---|---|
| verl | 0.9.0 | |
| Ray | 2.58.0 | |
| FlashInfer | 0.5.3 | |
| Triton Inference Server | 25.12 | |
| ROCm-RAG | 1.0.0 |
Detailed component changelogs#
verl 0.9.0#
This release updates verl in AMD LLM Extension from 0.7.1 to 0.9.0, and adds support for ROCm 10.0.0 on AMD Instinct™ MI300X, MI325X, MI350X, and MI355X GPUs with PyTorch 2.12.0 and Python 3.12 on Ubuntu 24.04.
Known issues#
Qwen2.5-Math-7B: set
max_position_embeddingsto32768after download.PYTORCH_ALLOC_CONF=expandable_segments:True: used to reduce OOM risk, and it can conflict with vLLM custom all-reduce. Default configs setvllm.disable_custom_all_reduce=Trueuntil that ROCm™ conflict is gone.SGLang:
attention_backendmust betriton.vLLM / SGLang ROCm™ fixes: some fixes are still applied in Dockerfiles rather than only in released wheels. The 10.0 recipe pins vLLM
v0.27.0instead ofmainbecause newer trees require atorch::stable::TensorAPI. SGLang is pinned to a commit validated on this Primus + vLLM ROCm™ stack (and patched in-tree forfused_add_rms_normand idempotent Qwen3-ASR config registration). Re-validateSGLANG_TAGif you bumpPRIMUS_TAGorVLLM_TAG.
Ray 2.58.0#
This release updates Ray in AMD LLM Extension from 2.55.1 to 2.58.0, and adds support for ROCm 10.0.0 on AMD Instinct™ MI300X, MI325X, MI350X, and MI355X GPUs with PyTorch 2.12.0 and Python 3.14 on Ubuntu 24.04.
Known issues#
RayTrain: it is recommended to use the legacy V1 API by explicitly setting
os.environ["RAY_TRAIN_V2_ENABLED"] = "0".