Command reference#

Examine#

rocm examine [--json] [--framework auto|pytorch|llama-cpp|skip]

Checks this computer’s GPU, ROCm install, engines, and managed setup folders — the command to run first to see whether a system is ready, and what rocm install sdk and rocm serve will see. --json emits a machine-readable report for diagnosis tooling instead of the human-readable summary. --framework controls which ML framework the --json report probes for its ROCm build and compiled GPU architectures: auto (the default) tries PyTorch, then falls back to llama.cpp; pytorch or llama-cpp probe only that framework; skip runs no framework probe at all, which is fastest and still enough to answer GPU and driver questions. --framework only affects the JSON report, not the human-readable one.

Diagnose and fix#

rocm diagnose [--symptom TEXT] [--top N] [--json] [--distro [NAME]]
rocm fix [<fix-id>] [--yes] [--dry-run] [--device-index N]

diagnose matches this machine against a fixed catalog of known ROCm/PyTorch/llama.cpp misconfigurations and ranks what it finds. It can only recognise failure modes that are in the catalog: no match means “not recognised”, not “nothing is wrong” — in that case it points you at where to report the symptom. Each result prints an id: and an apply with: command; the leading #1, #2 are ranking positions for reading order only — rocm fix takes the id, not the position.

  • --symptom takes raw error text to sharpen keyword scoring.

  • --top caps how many matches are shown in the human-readable output (default 5) — --json always emits the full, untruncated report.

  • --distro diagnoses a WSL distribution from the Windows host instead of this machine (nothing needs to be installed inside the distribution — name it only when more than one is installed). Inspecting remotely this way skips checks that need to read the distribution’s own environment (HSA_OVERRIDE_GFX_VERSION, PATH, the framework/ROCm pairing) — run rocm diagnose inside the distribution for those.

fix applies a known fix by the id: that diagnose reported — not the ranking position noted above, which isn’t a stable name. Run it with no id to list the whole catalog. Each fix is marked AUTO (this command carries out the change) or PRINT-ONLY (it prints the steps for you to run yourself — usually because the right command depends on a choice only you can make, sometimes because it also needs sudo or a reboot).

  • --dry-run shows any fix’s plan without changing anything.

  • --yes skips the interactive confirmation once you’ve reviewed it.

  • --device-index pins the discrete GPU index for fix-9-igpu-dgpu; without it, that fix only prints the rocminfo (Linux) or hipInfo.exe (Windows) query needed to find the index and makes no change, despite being marked AUTO.

ROCm installation#

rocm install sdk    [--channel release|nightly] [--format wheel|tarball]
                    [--version x.y.z | --build-date YYYY-MM-DD]
                    [--family gfx110X-all] [--prefix PATH] [--dry-run]
                    [--approve-replacing-active-default] [--yes]

rocm install driver [--dkms] [--yes] [--dry-run] [--reconcile]

rocm update         [--apply] [--runtime KEY] [--activate] [--dry-run]
                    [--json] [--timeout-secs SECS] [--yes]

install sdk downloads TheRock ROCm wheels into a Python environment managed by rocm-cli. An install with no active default runtime never prompts, but once a managed runtime is the active default every install sdk asks first, because the new install takes over as the active default. That gate is not scoped to the family or channel you are installing: a --family or --channel you have never installed before takes over the active default just as a same-family upgrade does, so it asks too. To approve that non-interactively — in scripts or CI, where the prompt would otherwise refuse — pass --approve-replacing-active-default, which is also what the refusal itself recommends and what ROCm CLI’s own non-interactive surfaces (chat, MCP, the dashboard) pass. --yes grants the same approval and approves installing required system packages (such as OpenMPI for vLLM), which means sudo; reach for it only where something can answer a sudo password prompt — which an unattended job cannot, unless it has passwordless sudo configured. In the default managed install root, the root and its manifest are keyed by version, so an upgrade or downgrade keeps the previous install on disk and only a same-version reinstall reuses the same root. --prefix opts out of that: the folder you name is used verbatim for every version, so successive installs into one prefix replace each other in place — and if the venv already there no longer runs its own Python, it is removed outright and rebuilt. The consent gate does not cover that: it asks about changing the active default runtime, not about what a named prefix loses. install driver installs the AMD kernel driver on Linux (DKMS or native package). update checks for a newer ROCm package; pass --apply to install it, or --dry-run to preview what --apply would do without changing anything (--dry-run does not require --apply). --runtime and --activate require --apply or --dry-run — pass one of those instead of naming a runtime or requesting activation on its own. --json prints the check result as a single line of JSON instead of text; --timeout-secs bounds its network calls (--timeout-secs requires --json; both --json and --timeout-secs conflict with --apply, and --json also conflicts with --dry-run). update --apply never prompts and needs no approval flag: selecting a runtime to update is itself the approval, and it leaves the active default alone unless you add --activate. update does accept --yes, for consistency with other mutating commands, but it grants nothing there — the approval line the update path prints never credits it.

ROCm 10 and newer ship from a different source layout. It is opt-in, and asking for it takes two things together: pin the version with --version, and name the exact GPU arch — the raw gfx code, not a family label:

rocm install sdk --version 10.0.0 --family gfx1200 --dry-run

A family label such as --family gfx120X-all is rejected for those versions rather than resolved to a guess, because the ROCm 10 packages publish one payload per exact arch and there is no bucket payload to fall back to. Run rocm examine to see the arch this machine reports.

For ROCm 10, install sdk asks uv to resolve Torch, torchvision, and torchaudio from their published dependency metadata, then validates that every selected framework package carries the same ROCm build identifier before it creates or changes a managed runtime.

Nothing about this happens on its own. Without a --version of 10 or newer, install sdk resolves the same release and nightly sources it always has, and it never quietly retries against the ROCm 10 sources when a lookup comes up empty — it tells you what it could not find instead.

Runtime management#

Manage multiple side-by-side ROCm runtimes:

rocm runtimes list
rocm runtimes activate <runtime-key>
rocm runtimes rollback
rocm runtimes uninstall <runtime-key> [--yes] [--dry-run]
rocm runtimes import <manifest-file> [--replace]
rocm runtimes adopt --python <path> [--root <path>] [--runtime-id ID]
                    [--runtime-key KEY] [--channel LABEL] [--replace]

uninstall prompts for confirmation unless --yes is passed; outside an interactive terminal --yes is required. --dry-run prints the plan and exits without prompting or making changes.

adopt registers an existing TheRock-based Python environment as a managed runtime. It does not work with standard ROCm package installs (for example, /opt/rocm); use rocm install sdk instead.

Disk space#

Each ROCm runtime keeps its own multi-gigabyte folder, so installing or updating a few times adds up. rocm storage shows where the space went and frees the parts that are safe to remove:

rocm storage [report] [--json]
rocm storage remove-old-installs [--keep N] [--dry-run] [--yes]
rocm storage remove-downloads [--dry-run] [--yes]

remove-old-installs keeps the two most recent installs for each channel, format, and GPU family, and never touches the install in use, the rollback target, or a folder rocm-cli did not create. “Most recent” means most recently installed rather than highest version, so after a deliberate downgrade the older version counts as the newer install. Because the count applies per channel, format, and GPU family, a machine that has tried several channels keeps --keep installs for each of them. Anything it declines to remove is listed with the reason, and --dry-run shows the whole plan without changing anything. remove-downloads clears cached archives that rocm-cli can download again; a cache folder that is a link to somewhere else is left alone rather than followed. The report also lists the uv package cache and downloaded models; those are shared with other tools and are never removed by rocm-cli.

Inference engines#

rocm engines list
rocm engines install <engine> [--runtime-id KEY] [--python-version X.Y] [--reinstall]
rocm engines shell <engine>   [--runtime-id KEY | --env-id ID] [--shell PATH]

Supported engines: lemonade, vllm.

Model serving#

Start a local OpenAI-compatible model server:

rocm serve <model> [--engine lemonade|vllm]
                   [--device gpu_required|gpu_preferred]
                   [--gpu auto|<index>]
                   [--runtime-id KEY | --env-id ID]
                   [--host HOST] [--port PORT]
                   [--verbose] [--foreground | --managed]
                   [--no-smoke-test]
                   [--allow-public-bind]
                   [--temperature TEMP] [--top-p PROB] [--max-tokens N]

--temperature (>= 0.0), --top-p (0.0-1.0), and --max-tokens (> 0) set server-wide sampling defaults for the launched engine. They apply only to vllm and lemonade; other engines reject them. For vLLM they are folded into a single --override-generation-config JSON object (--max-tokens maps to vLLM’s max_new_tokens); for Lemonade they pass straight through as llama.cpp’s --temperature, --top-p, and --n-predict flags. Each control is optional and independent — omit any of them to keep the engine’s own default.

rocm serve only reuses an already-running service for the same engine and model if its sampling controls (and other recipe settings) match the ones requested this time; otherwise it errors out instead of silently serving with different settings. If you previously started a service with --temperature (or another sampling flag) and now run rocm serve for the same model without flags — or with different ones — stop the existing service first (rocm services stop) or match the original flags.

By default the server runs in the background under rocm-cli’s supervision and prints a deployment summary — a progress indicator while it starts, then a table with the status, the full inference endpoint, the API-qualified model name, and a quick smoke test (time to first token and approximate tokens/sec). Control returns to your shell with the server still running; manage it later with rocm services (below).

--verbose (or --foreground) instead attaches to the server in the current terminal and streams every engine log line — use it to debug a startup problem. The server still runs as a managed background process, so while streaming you can press Ctrl-D to detach — the log stream stops, your shell comes back, and the server keeps running (manage it afterward with rocm services). Press Ctrl-C to stop the server instead. --managed is the explicit form of the default background behavior. --no-smoke-test skips the post-startup inference probe.

Which model form to pass depends on the engine your GPU selects. The Lemonade engine (Ryzen AI or Radeon) serves llama.cpp GGUF models — pass a GGUF repo with an explicit quantization variant, for example, rocm serve unsloth/Qwen3-0.6B-GGUF:Q4_0. The vLLM engine (Instinct) serves safetensors repos, such as rocm serve Qwen/Qwen2.5-1.5B-Instruct. A safetensors-only id has no GGUF build, so serving it through Lemonade fails rather than silently substituting a different model.

Some models (such as Llama) are gated and require HuggingFace authentication. Log in with huggingface-cli login or set HF_TOKEN in your environment before serving gated models.

--gpu selects which AMD GPU the server runs on. auto (the default) probes per-GPU VRAM with amd-smi and picks the lowest-numbered GPU that is idle and not already used by another rocm-cli server (managed or foreground), falling back to the GPU with the most free memory. Pass a single index (--gpu 1) to pin a specific device. The selected GPU is exposed to the engine via HIP_VISIBLE_DEVICES. Serving one model across multiple GPUs is not supported. Because selection uses the amd-smi ordinal but is applied via HIP_VISIBLE_DEVICES, rocm-cli warns when ROCR_VISIBLE_DEVICES is set, since the two orderings can diverge.

Manage background servers started with --managed:

rocm services list [--all]
rocm services logs <service-id>
rocm services stop <service-id> [--yes]
rocm services restart <service-id> [--yes]
rocm services remove <service-id> --yes
rocm services prune [--older-than-hours <n> | --any-age] [--dry-run] [--yes]

remove deletes one record that is no longer running, together with its log, its engine state file, and its endpoint key file; a running server is refused, so stop it first. prune does the same in bulk, always leaves running servers alone, and additionally clears leftover files whose record is already gone. Removal destroys both the log and the restart option for the records it takes, so prune only considers records untouched for 24 hours. Age is measured from when the record file was last written, so a stop, a restart, or a status correction all count as touching it. Pass --older-than-hours <n> for a different threshold, or --any-age to take every record that is not running however recent — that is the flag prune names in its own summary when it reports how many records it kept for being too recent. The two cannot be combined.

Dashboard#

rocm dash [--demo] [--replay <file>]

Full-screen TUI with Home, ROCm, Serving, Observe, and Chat tabs — GPU utilization graphs, active serving instances, benchmark results, guided actions,

and a chat tab backed by any configured provider. See Interactive interfaces for the tab breakdown.

  • --demo runs a deterministic synthetic session with no GPU or daemon needed, works on all platforms.

  • --replay <file> replays a recorded NDJSON session.

  • Live mode requires Unix domain sockets (Linux and WSL only).

Bench#

rocm bench load --endpoint URL [--model NAME] [--concurrency N,N,...]
                [--isl N] [--osl N] [--requests N] [--out FILE] [--auto-ramp]

Saturates a local OpenAI-compatible endpoint and reports rough client-side throughput — a local smoke test, not an official ROCm/AMD benchmark. load measures raw serving throughput with synthetic single-shot requests (the vLLM benchmark_serving shape); it does not reproduce agent-shaped, multi-turn, long-context tool traffic and isn’t comparable to *-agent-bench quality harnesses.

  • --endpoint is the OpenAI-compatible URL shown by rocm services list (a plain host address without /v1 also works); only http:// is accepted — https:// endpoints are rejected outright, since the load generator has no TLS backend compiled in.

  • --concurrency sweeps a comma-separated list of levels (default 1,8,32,64, each 1-128); --auto-ramp ignores --concurrency and instead ramps 1,2,4,8,16,32,64,128 automatically, stopping early once generation throughput plateaus or the request queue backs up.

  • --isl/--osl (input/output sequence length, default 1024 each) accept 1-32768, and --requests (default 128) accepts 1-10000.

  • Results are written to --out (default <data-dir>/bench/results.csv, where <data-dir> is ~/.rocm unless overridden), intended to match the path the daemon tails to feed the dashboard’s Observe tab. The CLI’s default output path and the daemon’s tailed path are computed independently, so if either the CLI’s data dir or the daemon’s bench_results_dir config has been customized, confirm they still point at the same file.

Chat#

rocm chat [--provider anthropic|openai|...] [--model NAME] [--prompt TEXT] [--tools]
          [--temperature TEMP] [--top-p PROB] [--max-tokens N]

Chat with an AI provider from the terminal. Reads from stdin when --prompt is omitted. --temperature, --top-p, and --max-tokens are optional sampling controls forwarded to the request; each is independent, so omit any of them to use the provider’s default.

ComfyUI#

Install and manage ComfyUI for image generation (alias: rocm comfy):

rocm comfyui install    [--runtime-id KEY] [--reinstall] [--dry-run] [--yes]
rocm comfyui start      [--host HOST] [--port PORT] [--no-open-browser] [--yes]
rocm comfyui stop       [--yes]
rocm comfyui status
rocm comfyui logs       [--lines N]
rocm comfyui models-path

None of install, start, or stop ever prompt for confirmation; --yes is accepted on each for consistency with other mutating commands but currently has no effect.

Automations#

rocm automations list
rocm automations enable <watcher-id>  [--mode observe|propose|contained]
rocm automations disable <watcher-id>

Optional background checks that can propose or apply changes automatically.

Configuration#

Show or change rocm-cli’s saved settings — the default engine and runtime, which runtime each engine prefers, local GPU telemetry opt-in, and the provider used for chat, automations, and ambiguous natural-language plans (including enabling providers and storing their API keys).

rocm config show
rocm config set-default-engine <engine>
rocm config clear-default-engine
rocm config set-default-runtime <runtime-id>
rocm config clear-default-runtime
rocm config set-engine <engine> [--runtime-id KEY | --env-id ID | --clear]
rocm config set-telemetry local|off
rocm config set-planner-provider <provider>
rocm config clear-planner-provider
rocm config enable-provider <provider>
rocm config disable-provider <provider>
rocm config set-provider-key <provider>
rocm config clear-provider-key <provider>

Setup#

rocm setup status
rocm setup reset

Manage first-time setup state. status shows whether first-time setup has completed; reset clears the recorded completed/dismissed state (nothing auto-triggers onboarding from this alone — open it manually from the dashboard, rocm dash: switch to the Observe tab, then press n). ROCm installs, API keys, and provider settings are left untouched.

Logs and cleanup#

rocm logs [--service <service-id>] [--search TERM ...]

rocm uninstall [--yes] [--dry-run]
               [--keep-binaries] [--keep-config] [--keep-data] [--keep-cache]

Shell completions#

rocm completions <shell> prints a completion script for the given shell to stdout. Supported shells are bash, zsh, fish, elvish, and powershell.

rocm completions <bash|zsh|fish|elvish|powershell>

Install the script for your shell:

# bash (per-user, no sudo; add this line to ~/.bashrc to persist)
source <(rocm completions bash)
# bash (system-wide; requires the bash-completion package)
rocm completions bash | sudo tee /etc/bash_completion.d/rocm > /dev/null

# zsh (per-user; the directory must be on $fpath and compinit must run)
mkdir -p ~/.zsh/completions
rocm completions zsh > ~/.zsh/completions/_rocm
# then in ~/.zshrc, before `compinit`:
#   fpath=(~/.zsh/completions $fpath)
#   autoload -Uz compinit && compinit

# fish
mkdir -p ~/.config/fish/completions
rocm completions fish > ~/.config/fish/completions/rocm.fish

# elvish (run once; re-running appends a duplicate block to rc.elv)
mkdir -p ~/.config/elvish
rocm completions elvish >> ~/.config/elvish/rc.elv

# powershell (current session only; to persist, append the output to $PROFILE)
rocm completions powershell | Out-String | Invoke-Expression