vLLM adapter#
rocm-engine-vllm is a first-party adapter around an existing vLLM
installation. It is intended for Linux and WSL ROCm GPU serving.
The adapter does not install vLLM automatically and does not run CPU mode.
Install or build vLLM in a ROCm-capable Python environment first, then make the
vllm command visible to ROCm CLI.
Native Windows vLLM serving is skipped in this adapter. Use WSL or Linux for vLLM ROCm serving, or choose a different engine explicitly. No CPU fallback is used.
For ROCm CLI-managed TheRock runtimes, prefer building vLLM from source against the existing TheRock PyTorch stack. A prebuilt vLLM ROCm wheel can replace the TheRock torch packages or target a different ROCm soname set; that is not a valid no-fallback setup for ROCm CLI GPU serving.
Building from source compiles HIP sources, so it needs the ROCm compiler
toolchain. That is opt-in: install the runtime with rocm install sdk --devel,
or the build fails on a missing compiler. A runtime installed without --devel
can still serve an already-built vLLM.
Torch alignment on engine install#
Installing an engine into a managed TheRock runtime can change the torch in that
runtime. Two installers write torch into the same environment: the SDK install
writes TheRock’s build, and the engine install then writes the build from its own
index. So rocm engines install settles which one stays and prints the result
as a torch_alignment: line.
A torch that already executes a GPU kernel against the installed SDK is kept
exactly as it is, whichever installer put it there. Otherwise the runtime is
moved to the SDK’s build of the torch release the engine pins: the release
comes from the engine, which was built against it, and the build comes from the
SDK, whose libraries it has to load. A device_check: line reports what the
result can actually do, and a realignment also reports what the runtime could do
before it.
Set ROCM_CLI_DISABLE_TORCH_ALIGNMENT to keep whatever torch is installed and
skip the replacement:
ROCM_CLI_DISABLE_TORCH_ALIGNMENT=1 rocm engines install vllm --yes
Any value works, including an empty one: the variable being set is the signal.
The install then reports torch_alignment: disabled, naming both the build it
would have installed and the one it kept. The device check still runs, so an
opt-out that leaves the runtime unable to serve says so rather than failing later
during serving.
Use it when you are deliberately running a torch the alignment would replace (a locally built wheel, a version under test, or a stack pinned for a reproduction). It is an escape hatch, not a supported configuration: the resulting combination is not validated against the supported matrix, and a runtime that cannot execute a kernel will fail at serving time.
ROCm 10.x wheel discovery#
For most ROCm SDK versions, rocm engines install vllm pins a fixed vLLM wheel
and index URL. Any ROCm SDK 10.x version is different: AMD publishes vLLM,
flash-attn, and amd-aiter there under a rotating dev-tag filename (for example
vllm-0.27.1.dev5+rocm10.0.0.gf46a9dfe2.d20260826-cp314-cp314-linux_x86_64.whl),
so there is no fixed filename to pin in the adapter. The wheel’s own tag shows
the other thing that changes on this route: ROCm 10.x’s vLLM, flash-attn, and
amd-aiter are published for cp314 only, unlike every earlier ROCm SDK
version’s cp312 wheels. ROCm CLI provisions a matching interpreter for a
10.x install automatically; a Python already on PATH or already installed
as ROCm CLI’s managed Python is rejected if it is not cp314, with a message
naming the required tag.
This route is selected by major.minor, so 10.0.0 and 10.1.0 discover
through different rows: they use different index URLs and different vLLM
minors, because AMD stages ROCm 10.1’s frameworks on a separate host from
10.0’s production index. Patch and any dev/pre-release suffix are still
ignored within a row, since AMD rotates those constantly.
The install resolves each package’s current wheel, including torch, from the
row’s index with uv pip install --dry-run --reinstall, parses the version it
reports it would install, then reinstalls pinned to that exact version.
tensorizer is not discovered or pinned this way: it has no ROCm-specific
build, and vllm’s own wheel metadata already declares an exact tensorizer
dependency, so it is left to vllm’s own dependency resolution rather than
risk a conflicting pin of its own. Every resolved pin that carries a
+rocmX.Y local version is checked against the SDK’s own major.minor before
installing, so an index that happens to serve more than one ROCm line at once
cannot silently install the wrong line’s wheel onto this SDK. If AMD’s index
has no compatible build for a package, the resolver fails and the install
fails rather than falling back to an unpinned or CPU install. The final install
of vllm/flash-attn/amd-aiter also resolves vllm’s own plain-PyPI transitive
dependencies (e.g. lm-format-enforcer), which AMD’s index doesn’t host; a
generated uv.toml sets ignore-error-codes = [403] for that index so uv
falls through to PyPI for those instead of treating the index’s 403 as fatal.
That same full-dependency resolve can also pull in an unconstrained torch
from PyPI, undoing the exact ROCm pin just installed, and transitively an
unconstrained torchvision/torchaudio with it; the install re-pins torch,
plus whatever of torchvision/torchaudio was already installed by the SDK
install, back to their prior exact builds immediately afterwards.
Every other
ROCm SDK version, including 7.2.3, keeps using the static pin table; an SDK
version with no matching row there falls back to the table’s default pin,
unless its major release matches a discovery-table entry, in which case
guessing the default pin would very likely install an ABI-incompatible build,
so the install fails closed instead with a message naming the detected
version and pointing at ROCM_CLI_VLLM_ROCM_INDEX_URL as the way to install
anyway.
Discovery paths and checks#
Supported discovery paths:
ROCM_CLI_VLLM_COMMAND=path_to_vllm, wherepath_to_vllmis the absolute path to thevllmexecutableROCM_CLI_VLLM_PYTHON=path_to_python, wherepath_to_pythonis the absolute path to a Python interpreter that has a siblingvllmcommandthe active ROCm CLI-managed TheRock runtime, if vLLM has been installed into that Python environment
vllmonPATH
Useful checks:
rocm-engine-vllm detect
rocm-engine-vllm capabilities
rocm-engine-vllm resolve-model Qwen/Qwen3.5-4B --device-policy gpu_required
python scripts/vllm_therock_gpu_test.py --self-test
GPU acceptance check#
Run the acceptance script against a built adapter:
python3 scripts/vllm_therock_gpu_test.py \
--engine target/debug/rocm-engine-vllm \
--model facebook/opt-125m
The acceptance script is Linux or WSL only. It requires vLLM to be discoverable
through a ROCm CLI-managed TheRock runtime manifest, launches with
gpu_required, checks /health and /v1/completions, and verifies loaded
ROCm libraries come from the managed TheRock SDK wheel directories. It rejects
external vLLM command overrides and does not allow CPU fallback. It defaults
to the active exact runtime key; if --runtime-id is passed, use an exact
runtime key or an unambiguous runtime id.
Source build notes#
On WSL, the tested source build needed vLLM ROCm platform detection to use
TheRock PyTorch device data when amdsmi is unavailable, and needed vLLM’s
ROCm GPTQ half-atomic compatibility path enabled for TheRock 7.13 headers.
On the MI300X/gfx942 TheRock 7.13 runtime, current vLLM source required the
GPTQ compatibility guard in
csrc/libtorch_stable/quantization/gptq/compat.cuh to include HIP 7.13:
- (defined(USE_ROCM) && (HIP_VERSION_MAJOR * 100 + HIP_VERSION_MINOR) < 713)
+ (defined(USE_ROCM) && (HIP_VERSION_MAJOR * 100 + HIP_VERSION_MINOR) <= 713)
Without that patch, q_gemm.hip fails to compile because TheRock 7.13 headers
do not expose the half/half2 atomicAdd overloads used by vLLM’s GPTQ
kernel. With the patch, the live acceptance harness passed on
facebook/opt-125m and verified HIP/BLAS libraries loaded from the managed
TheRock SDK wheel directories.
Serve a model#
Serve a model through ROCm CLI:
rocm serve Qwen/Qwen3.5-4B --engine vllm --device gpu_required --managed
GPU selection#
Use --gpu to choose the AMD GPU vLLM runs on:
# Default: first free GPU (auto)
rocm serve Qwen/Qwen3.5-4B --engine vllm --managed
# Pin a specific GPU
rocm serve Qwen/Qwen3.5-4B --engine vllm --gpu 1 --managed
ROCm CLI pins the device via HIP_VISIBLE_DEVICES. Serving one model across
multiple GPUs is not supported.
GPU memory#
vLLM claims a fixed fraction of each GPU’s total VRAM (not of the free VRAM, and not scaled to the model) for weights plus KV cache. On a large card a small model therefore still reserves a large slice.
ROCm CLI sets no --gpu-memory-utilization of its own, so vLLM’s own default
applies unless a value comes from somewhere else: either a model’s catalog
recipe or, taking precedence over it, the flag below:
rocm serve <model> --engine vllm --gpu-memory-utilization 0.3 --managed
The value is a fraction in (0, 1] of total device VRAM. Lower it to leave room
for a display, another workload, or a second server; raise it to give a large
model more KV cache. Applies to vLLM only; it is ignored, with a note in the
serve output, for other engines. An out-of-range or unparsable value fails the
command rather than falling back silently.
Earlier releases pinned this to 0.80 to leave display/WSL headroom. That pin is
gone, so an unchanged command now reserves vLLM’s own (higher) default. Pass
--gpu-memory-utilization 0.8 to restore the previous reservation.
Tool calling#
The TUI chat tab attaches tool definitions to every chat request. vLLM rejects
those with HTTP 400 unless it is launched with --enable-auto-tool-choice and
a matching --tool-call-parser. vLLM does not auto-detect the parser and it is
model-specific, so ROCm CLI never guesses one:
Built-in catalog models carry the correct parser in their recipe metadata, so tool calling works out of the box (for example, Qwen family →
hermes, Llama 3 →llama3_json).Other models (arbitrary Hugging Face repos, or a catalog model forced onto vLLM without authored metadata) need an explicit parser:
rocm serve <model> --engine vllm --tool-call-parser hermes --managed
--tool-call-parserimplies--enable-auto-tool-choice, overrides any catalog default, and applies to vLLM only. Common values:hermes,llama3_json,mistral. Without it, plain chat still works but tool calls return HTTP 400.