llama.cpp inference on ROCm#

llama.cpp is an open-source inference engine for running LLMs locally. It runs quantized models efficiently on everyday hardware and has become a pillar of local LLM inference. This page explains how to set up and run llama.cpp on AMD Radeon™ GPUs and Ryzen™ APUs.

Device family
Operating system
ROCm version
ROCm installation

See also

Prerequisites#

To run llama.cpp with ROCm on Linux, you need the following prerequisites:

  • GPU platform: AMD Instinct accelerators, and AMD Radeon discrete GPUs and Ryzen APUs supported by ROCm 7.14.0. For the list of supported devices, see the ROCm compatibility matrix.

Install ROCm#

llama.cpp requires a ROCm installation on the host system.

Install llama.cpp#

AMD does not publish its own binaries for llama.cpp. Download a prebuilt ROCm-enabled release from the llama.cpp releases page. Release assets are named in the following format:

Note

The archive contains only the llama.cpp executables and their ggml backend libraries. It does not include the ROCm runtime, which must be installed separately as described in the previous section.

Verify the installation#

Run a llama.cpp example#

Once your llama.cpp environment is set up, experiment with the following steps to run a model and benchmark your installation. The prebuilt llama.cpp release includes the executables to exercise the functionality of your installation.

The two most popular use cases are:

  • llama-cli: The main executable to run the model interactively or get a response to a prompt.

  • llama-bench: Run a benchmark of your model with different configurations.

See also

See Models (llama.cpp docs) for a list of GGUF-formatted models available to download from Hugging Face.

llama-cli#

  1. Use the CLI tool to start the client, replacing <model>.gguf with the path to a model on your system:

  2. A prompt appears when the client is ready, and you can start interacting with the model:

    > Explain what a GPU kernel is in one sentence.
    A GPU kernel is a specialized program within a GPU (Graphics Processing Unit) designed to execute
    instructions for rendering graphical content, such as images or video, on a GPU's hardware architecture.
    
    [ Prompt: 593.4 t/s | Generation: 286.4 t/s ]
    
  3. To exit, enter /exit or press Ctrl+C.

To send a single prompt and exit instead of starting an interactive session, add the -p and -st flags:

Note

llama-bench#

  1. Use the CLI tool to start the application, replacing <model>.gguf with the path to a model on your system:

  2. The result of the command above should be similar to the following when running on an AMD Radeon RX 7900 XT system:

    ggml_cuda_init: found 1 ROCm devices (Total VRAM: 20464 MiB):
      Device 0: AMD Radeon RX 7900 XT, gfx1100 (0x1100), VMM: no, Wave Size: 32, VRAM: 20464 MiB
    load_backend: loaded ROCm backend from C:\llamacpp\ggml-hip.dll
    load_backend: loaded RPC backend from C:\llamacpp\ggml-rpc.dll
    load_backend: loaded CPU backend from C:\llamacpp\ggml-cpu-haswell.dll
    | model                          |       size |     params | backend    | ngl |            test |                  t/s |
    | ------------------------------ | ---------: | ---------: | ---------- | --: | --------------: | -------------------: |
    | qwen2 1B Q4_K - Medium         | 462.96 MiB |   630.17 M | ROCm       | 999 |            pp16 |     2406.95 ± 704.75 |
    | qwen2 1B Q4_K - Medium         | 462.96 MiB |   630.17 M | ROCm       | 999 |            pp32 |     3842.56 ± 500.15 |
    | qwen2 1B Q4_K - Medium         | 462.96 MiB |   630.17 M | ROCm       | 999 |            pp64 |    7104.29 ± 1218.73 |
    | qwen2 1B Q4_K - Medium         | 462.96 MiB |   630.17 M | ROCm       | 999 |           pp128 |     8628.19 ± 928.69 |
    | qwen2 1B Q4_K - Medium         | 462.96 MiB |   630.17 M | ROCm       | 999 |           pp256 |   16444.29 ± 1662.23 |
    | qwen2 1B Q4_K - Medium         | 462.96 MiB |   630.17 M | ROCm       | 999 |           pp512 |   24628.13 ± 2498.10 |
    | qwen2 1B Q4_K - Medium         | 462.96 MiB |   630.17 M | ROCm       | 999 |          pp1024 |     26143.12 ± 73.58 |
    | qwen2 1B Q4_K - Medium         | 462.96 MiB |   630.17 M | ROCm       | 999 |            tg64 |        373.29 ± 9.81 |
    | qwen2 1B Q4_K - Medium         | 462.96 MiB |   630.17 M | ROCm       | 999 |           tg128 |        375.70 ± 3.49 |
    
    build: 78ec4c378 (10539)
    
  3. The backend column shows ROCm when the GPU is in use. If it shows CPU, the ROCm libraries were not found. Review the environment configuration steps to troubleshoot.

Flash Attention is enabled with -fa 1 and is supported on Radeon GPUs through rocWMMA: