serving-llms-on-instinct#
The goal of this skill is to teach your AI agent to bring up a vLLM OpenAI-compatible endpoint on an AMD Instinct™ GPU host on ROCm: detecting the GPU, validating the environment, picking the right vLLM recipe for the model, checking the model fits VRAM, launching the container, and verifying the endpoint responds.
What you’ll end up with: a running vLLM endpoint on your Instinct box (in a Docker
container built from the ROCm vLLM image), sized to the GPUs you actually have, and
ready to answer OpenAI requests via /v1/chat/completions.
Prerequisites#
A supported AMD Instinct™ GPU: MI300X / MI325X / MI300A (
gfx942) or MI350X / MI355X (gfx950).detect.pyreportsgfx_version; anything else is out of scope. This skill explicitly does not cover MI250X, MI100, consumer Radeon (RX series), or Ryzen AI / NPU.ROCm driver and
amd-smiinstalled on the host, with/dev/kfdand/dev/dripresent. Check the driver is loaded withlsmod | grep amdgpu.Docker running and accessible —
docker psmust work withoutsudo. If it errors on/dev/kfdpermissions, add yourself to the GPU groups:sudo usermod -aG video,render $USER(requires re-login).Disk space in
~/.cache/huggingfacefor the model weights (the default model below is roughly 18 GB at BF16).A HuggingFace token in
HF_TOKENonly for gated models (Llama, Gemma). The default model (Qwen3.5) needs none. For gated models the token must belong to an account that has accepted the license athuggingface.co/<model_id>— a valid token without license acceptance fails with an opaqueEngine core initialization failed.Node.js ≥ 18, required by the
skillsCLI used in Step 2 (npx skills ...). Check withnode -v.No Instinct hardware handy? AMD Developer Cloud offers hosted MI300X instances that satisfy the requirements above. Note this is just a way to get a host — the skill itself does not provision or onboard cloud instances; it expects a machine that’s already up.
Step 1 - Understanding which skills are available#
Start in a clean scratch directory, then run
claude "Which skills can you see?" --model sonnet. You should see a list of skills that does not include anything about serving LLMs on Instinct / AMD GPUs.Confirm the scratch directory has no
AGENTS.md,CLAUDE.md,.claude/skills, or.agents/skills. Existing agent instructions or installed skill copies can change discovery and invalidate the before/after comparison. Do not delete instructions from a real project; use a clean scratch directory instead.
Step 2 - Enabling claude to see serving-llms-on-instinct#
Install the skill with the
skillsCLI. Run this in your terminal, not inside Claude:
npx skills add amd/skills --skill serving-llms-on-instinct --agent claude-code
Run
claude "Which skills can you see?". You should see a list of skills that now includesserving-llms-on-instinct.
Step 3 - Running the skill#
You can run this skill directly on the Instinct host if Claude is installed there. Otherwise if the host is reachable over SSH, provide its server address so Claude can connect to it and run the skill remotely.
Run claude on your Instinct host with this prompt:
Serve Qwen/Qwen3.5-9B on this AMD Instinct GPU with vLLM.
Claude should:
Detect the GPU: read
gfx_version,vram_gb,gpu_count, androcm_version, and map the gfx target to the hardware (gfx942→ MI300X/MI325X/MI300A,gfx950→ MI350X/MI355X). The gfx version drives precision support and the workarounds applied later.Validate the environment: check
/dev/kfd,/dev/dri, Docker accessibility, NUMA balancing, hipBLASLt, andHF_TOKEN, classifying each as error / warning / advisory. Safe fixes (like disabling NUMA balancing, which causes latency spikes for GPU workloads) are applied with--auto-fix. It does not proceed if the environment isn’t ready.Refresh the vLLM recipe cache if it’s older than 24 hours, pulling the latest model recipes from vllm-project/recipes and resolving the current ROCm vLLM Docker tag (~10 seconds). A stale cache still works if the refresh fails.
Check the model is actually servable: reject blacklisted models — diffusion/image and audio generation, embeddings, rerankers, ASR models needing an audio pipeline, and models that require an unreleased vLLM nightly — and suggest an alternative rather than failing at launch.
Build the config from the model’s recipe: mandatory AMD Docker flags, the HF cache mount, gfx-specific env defaults, the recipe’s vLLM args plus tool-calling and reasoning parsers, and the ROCm image (
vllm/vllm-openai-rocm, never the CUDA-onlyvllm/vllm-openai). It picks a precision variant the GPU supports — ongfx942FP8 is FNUZ and MXFP4 compute is emulated; ongfx950MXFP4 is native; NVFP4 is rejected on both.Check it fits VRAM: estimate weight memory and KV cache from the HuggingFace Hub API (no download), reserving ~4 GB for vLLM runtime overhead, then decide from what’s left — cap
--max-model-lenif context is KV-limited, raise tensor parallelism if the weights need more than one GPU, or switch to a quantized checkpoint (recipe variant, same-provider FP8, or anamd/Quark model) if they don’t fit at all.Confirm the plan with you: present a summary (model, precision and why, weight memory, GPU, TP degree, achievable context, port) and wait for you to approve before launching. If it swapped in a quantized alternative, it says so and explains why.
Launch and verify: check the port is free, run the container, poll
/healthuntil the server answers (a 503 while loading is normal), then send a warmup request — the first inference compiles HIP kernels and takes ~40-45 seconds ongfx942— and hand back a connection table.
Note the one place this skill does retry: if the model fits but vLLM OOMs during HIP graph
capture, it relaunches once with --enforce-eager, which frees 1-2 GB at the cost of slightly
higher decode latency. That is a bounded, known-cause retry, not a debugging loop.
Step 4 - Talk to the endpoint#
Once Claude reports that the endpoint is healthy, follow the terminal instructions it provides to verify and use it. The exact curl command may vary depending on the model and serving configuration, so do not run the example below; use the command provided by Claude instead
curl -s http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"Qwen/Qwen3.5-9B","messages":[{"role":"user","content":"Hello"}],"max_tokens":128}'
To stop the endpoint: docker rm -f <container name from the connection table>.
Step 5 - (Optional) Going beyond#
A bigger model: ask for something that needs more than one GPU, e.g. “Serve Qwen/Qwen3-235B-A22B”. Claude re-runs the fit check, raises tensor parallelism, and adds
--distributed-executor-backend mpfor MoE models on multi-GPU.A model that doesn’t fit: ask for a large dense model on a single GPU and watch it find a quantized alternative (recipe FP8 variant, a same-provider FP8 checkpoint, or an
amd/Quark model) instead of launching into an OOM.Gated models:
export HF_TOKEN=..., accept the license on HuggingFace, then ask formeta-llama/Llama-3.2-....Longer context: if Claude reports the context is KV-limited, ask about FP8 KV cache (
--kv-cache-dtype fp8) to buy more sequence length out of the same VRAM.Share a host: on a busy multi-GPU node, tell it which GPUs to use (“only use GPUs 0 and 1”) — it restricts with
HIP_VISIBLE_DEVICES.Drive it remotely: every script accepts
--host user@hostname(or theROCM_SSH_HOST/ROCM_SSH_USERenv vars) and runs over SSH, so you can serve on a remote Instinct node from your laptop. Key-based SSH must already work.
Step 6 - (Optional) Try to get things done without AMD Skills#
Remove the added skill and rerun the experiment above. The skills CLI installs a copy
under both .claude/skills/serving-llms-on-instinct and
.agents/skills/serving-llms-on-instinct, so delete both (otherwise the leftover copy
keeps the skill active and the comparison isn’t clean). Without the skill, common issues
include:
Reaching for
vllm/vllm-openai— the CUDA-only image — instead ofvllm/vllm-openai-rocm, so nothing finds a GPU.Missing the mandatory AMD Docker flags (
--device /dev/kfd,--device /dev/dri,--group-add video/render,--cap-add SYS_PTRACE,--security-opt seccomp=unconfined,--ipc=host), so the container starts but sees no GPU or dies during ROCm JIT.No VRAM fit check: launching a model that doesn’t fit and looping on
Engine core initialization failed, or fitting the weights but OOM-ing during HIP graph capture without knowing--enforce-eageris the fix.Picking an NVFP4 checkpoint, which has no dequant kernel on ROCm and will not load — or assuming MXFP4 is native on MI300X, where it is emulated.
Missing the gfx-specific workarounds, most notably
VLLM_ROCM_USE_AITER_FP4BMM=0ongfx942, which otherwise segfaults during warmup.Missing per-model args like
--block-size 1for MLA models (DeepSeek-R1/V3, Kimi-K2.5), which silently falls back to a slower attention path.Trying to serve a diffusion, embedding, or ASR model as a chat endpoint instead of catching it up front.
Declaring success at
/health200 without a warmup request, so your first real call appears to hang for ~40 seconds of HIP kernel compilation.Providing a knowledge article instead of actually bringing up a working endpoint.