SGLang inference and serving on ROCm#
SGLang is an open-source library for fast, memory-efficient LLM inference and serving. This page describes how to set up and run SGLang on AMD GPUs using either a prebuilt Docker image (recommended) or pip. It applies to supported AMD GPUs and platforms.
ROCm version
10.0.0
7.14.1
7.14.0
Device family
AMD Instinct™
AMD Radeon™
SGLang version
0.5.15
0.5.13
Installation method
Docker
Prerequisites#
Ensure the host system has Docker Engine installed. For more guidance on running ROCm workloads in Docker containers, see Run ROCm Docker containers.