AMD LLM Extension 26.09 release notes#

3 min read time

Applies to Linux

This is the seventh release of the AMD LLM Extension toolkit, an open-source software toolkit built on the ROCm platform for large language model (LLM) extensions, integrations, and performance enablement on AMD GPUs. The domain brings together training, post-training, inference, and orchestration components to make modern LLM stacks practical and reproducible on AMD hardware.

Release highlights#

Note

AMD LLM Extension 26.09 introduces updates to two components (verl and Ray) as part of the toolkit; other components remain unchanged (FlashInfer, ROCm-RAG, and Triton Inference Server).

This release updates the following components with support for ROCm 10.0.0:

  • verl (Volcano Engine Reinforcement Learning for LLMs) is an open-source framework for reinforcement-learning post-training of large language models and is the reference implementation of HybridFlow.

  • Ray is an open-source framework for scaling Python and AI workloads, providing the distributed compute and orchestration layer used by verl for multi-GPU and multi-node training.

System requirements#

AMD LLM Extension components span a range of ROCm version requirements depending on the specific extension. Ensure you follow the installation instructions for each component, which list the exact ROCm dependencies, or refer to the Compatibility matrix to verify the supported ROCm versions.

AMD LLM Extension components#

The following table lists AMD LLM Extension component versions for the 26.09 release. Click to go to the component’s source on GitHub.

Name Version Source
verl 0.9.0
Ray 2.58.0
FlashInfer 0.5.3
Triton Inference Server 25.12
ROCm-RAG 1.0.0

Detailed component changelogs#

verl 0.9.0#

This release updates verl in AMD LLM Extension from 0.7.1 to 0.9.0, and adds support for ROCm 10.0.0 on AMD Instinct™ MI300X, MI325X, MI350X, and MI355X GPUs with PyTorch 2.12.0 and Python 3.12 on Ubuntu 24.04.

Known issues#

  • Qwen2.5-Math-7B: set max_position_embeddings to 32768 after download.

  • PYTORCH_ALLOC_CONF=expandable_segments:True: used to reduce OOM risk, and it can conflict with vLLM custom all-reduce. Default configs set vllm.disable_custom_all_reduce=True until that ROCm™ conflict is gone.

  • SGLang: attention_backend must be triton.

  • vLLM / SGLang ROCm™ fixes: some fixes are still applied in Dockerfiles rather than only in released wheels. The 10.0 recipe pins vLLM v0.27.0 instead of main because newer trees require a torch::stable::Tensor API. SGLang is pinned to a commit validated on this Primus + vLLM ROCm™ stack (and patched in-tree for fused_add_rms_norm and idempotent Qwen3-ASR config registration). Re-validate SGLANG_TAG if you bump PRIMUS_TAG or VLLM_TAG.

Ray 2.58.0#

This release updates Ray in AMD LLM Extension from 2.55.1 to 2.58.0, and adds support for ROCm 10.0.0 on AMD Instinct™ MI300X, MI325X, MI350X, and MI355X GPUs with PyTorch 2.12.0 and Python 3.14 on Ubuntu 24.04.

Known issues#

  • RayTrain: it is recommended to use the legacy V1 API by explicitly setting os.environ["RAY_TRAIN_V2_ENABLED"] = "0".