Kimi-K3 (optimized build)#
Kimi-K3 on MI355X using johnqin2025/kimi-k3-dspark — an image carrying an FP8
pre-route / shared-expert kernel cluster, an FP8 latent-MoE tail, a fused MLA
output gate and a tri-projection dispatch — plus an optional DSpark draft model
for speculative decoding.
A separate entry from Kimi-K3 rather than a combination on it, because
it is a different artifact. Same vLLM commit (g5f76ae224) as the stock image, so
the delta is kernels, not engine version.
Every manifest here is a template — fill in the placeholders first
A bare kubectl apply on these files is accepted by the API server (the
placeholder passes CRD validation) and then fails minutes later at mount time.
Placeholder |
In |
How to find it |
|---|---|---|
|
Mixed, Mixed + DSpark |
site-specific, no default — see below |
|
Mixed, Mixed + DSpark |
a node with 8 free GPUs that has the weights at that path |
|
PD, PD + DSpark |
two nodes on a routable RoCE fabric |
|
PD, PD + DSpark |
per node — they need not be equal |
The model path is site-specific and a mixed fleet may not agree with itself. On
the fleet these were validated on, the two GPU nodes use different paths for the
same weights. Candidates usually look identical (same shard count, same layout)
while only one is local storage, so confirm the one you pick by mounting that
directory in a throwaway pod on the node and reading df -PT — see the repo
README for the commands. Mounting a parent instead reports the root filesystem and
will call a network-mounted weights directory local.
Leaving a placeholder in has two different outcomes, neither obvious. <MODEL_DIR>
creates a Pod that fails at mount time. <NODE> creates no Pod at all —
apply prints created, get pods shows nothing, and the only error is in the
operator’s log in the infera-system namespace.
<NODE> exists because both the server and the worker mount the weights by
hostPath and are scheduled independently — without it they can land on different
nodes.
1. Choose the combination#
8 GPUs on one node, no speculative decoding. The default, and the best tokens per GPU of the four.
sed -e 's|<MODEL_DIR>|/mnt/local-nvme/models|' -e 's|<NODE>|node-a|' \
examples/recipes/kimi-k3-optimized/aggregated/deploy.yaml | kubectl apply -f -
Placeholder |
What to put there |
|---|---|
|
Directory on that node containing |
|
Hostname of the node ( |
Start here unless you have a specific reason not to. The other three trade GPUs or added complexity for latency, and one of them needs a second node.
Same 8 GPUs, plus speculative decoding: a block-diffusion draft model that produces 7 tokens in one parallel pass, verified against the target.
sed -e 's|<MODEL_DIR>|/mnt/local-nvme/models|' -e 's|<NODE>|node-a|' \
examples/recipes/kimi-k3-optimized/aggregated-dspark/deploy.yaml | kubectl apply -f -
Same two placeholders as Mixed, but <MODEL_DIR> must also contain
Kimi-K3-DSpark/ alongside Kimi-K3/.
The draft is a community checkpoint,
Inferact/Kimi-K3-DSpark — there
is no official Moonshot draft for Kimi-K3 — and must be downloaded alongside the
target into the same <MODEL_DIR>.
Speculation earns its keep at low concurrency, where there is idle capacity to spend on drafting. As concurrency rises the verify batch grows, acceptance falls, and the advantage decays; at high concurrency plain Mixed is the better choice. Where that crossover sits depends on your workload — measure it on yours rather than taking a number from here.
The old concurrency ceiling is gone
On the …-20260801 image this died at c>=16 with AssertionError: AiterMLA flattened verify requires a uniform decode query len. The image pinned here fixes
it — c=16/32/64 all complete with 0 restarts and the assertion never appears — and
speculation is genuinely still on (num_spec_tokens=7, CUDA-graph captured,
running the draft eagerly count 0), so this is a fix rather than speculation
being quietly disabled.
If you pin the older digest, the old ceiling still applies to you.
Prefill and decode on separate nodes, 8 GPUs each, KV handed between them over RDMA.
sed -e 's|<PREFILL_NODE>|node-a|' -e 's|<DECODE_NODE>|node-b|' \
-e 's|<PREFILL_MODEL_DIR>|/mnt/local-nvme/models|' \
-e 's|<DECODE_MODEL_DIR>|/mnt/array/models|' \
examples/recipes/kimi-k3-optimized/disaggregated/deploy.yaml | kubectl apply -f -
Placeholder |
What to put there |
|---|---|
|
Hostname of the node that runs prefill ( |
|
Hostname of the node that runs decode — it receives that KV over RDMA and generates the tokens. Must be a different node, and the two must be able to reach each other over the RoCE fabric. |
|
Directory on the prefill node containing |
|
The same, on the decode node. Frequently a different path — the nodes do not have to agree, and on a mixed fleet they often do not. |
Both directories are hostPath mounts, so each node reads its own local copy;
there is no shared volume. Neither path is discoverable from this page — find each
one on its own node, and confirm it is local storage rather than a network mount,
before substituting.
Twice the hardware, and it never wins per GPU
PD does not improve tokens per GPU at any concurrency measured here. What it buys is headroom past what a single node can serve, and lower latency at high concurrency, because decode is no longer interleaved with prefill on the same GPUs. Below that it is a straight loss on both counts — use Aggregated.
Requires both nodes on a mutually routable RoCE fabric: the KV handoff is RDMA and there is no TCP fallback. Each node reads its own local copy of the weights, and the two paths need not be the same.
Checking the fabric before you deploy
Use Preflight — it covers RDMA device and link state, cross-node RoCE bandwidth, Mooncake KV-transfer bandwidth measured separately over RDMA and TCP, and whether the KV path is on local NVMe:
# one node, all image-independent checks
python -m infera.tools.preflight --dump-path output/preflight
# just the network probes
python -m infera.tools.preflight --network
# both nodes at once, under SLURM, rendering one combined report
NODES=<node-a>,<node-b> PARTITION=<partition> IMAGE=<image> \
infera/tools/preflight/run_preflight_slurm.sh
It writes <dump-path>/<host>.json per node plus a combined HTML report. GPU perf
and ais-check only run inside the engine container; run it there for those.
See Preflight for the full check list, thresholds and the
multi-node path.
The Mooncake rows are the ones that matter for this recipe: they report KV-move
bandwidth over rdma and over tcp separately, so a fabric that will silently
serve at TCP speed shows up as a number rather than as a slow deployment.
A second, automatic check runs at launch
infera/common/disagg_preflight.py validates the disaggregated config before the
engine subprocess starts, and fails fast rather than hanging. It catches a worker
advertising a non-routable host (0.0.0.0, 127.0.0.1) to etcd — the peer then
cannot reach its bootstrap endpoint across nodes — and configurations prone to
silent TCP fallback. It is pure config validation, so it cannot tell you the NIC
itself is healthy; that is what the tool above is for.
If you only want to know whether the container can see the RDMA devices, note that
ibv_devices is not installed in these images — reading its not found as “no
RDMA devices” produced two false TCP-fallback diagnoses during this work. Ask the
library, in a prefill or decode pod (the router pod is not privileged and does
not mount /dev/infiniband, so it reports 0 and reproduces that same false
negative):
POD=$(kubectl -n infera get pod -o name \
-l infera.amd.com/deployment=kimi-k3-opt-pd,infera.amd.com/service=decode | head -1)
kubectl -n infera exec $POD -c main -- python3 -c '
import ctypes; lib = ctypes.CDLL("libibverbs.so.1")
lib.ibv_get_device_list.restype = ctypes.POINTER(ctypes.c_void_p)
n = ctypes.c_int(0); lib.ibv_get_device_list(ctypes.byref(n)); print(n.value)'
For reachability rather than device visibility, see
the RoCE note on the recipes index — on a routed L3 fabric an unbound
ping6 picks the wrong source rail and reports “No route”, which reads as “these
nodes cannot do PD” when they can.
the disaggregated topology plus speculation. The lowest latency of the four — the decode role is not competing with prefill for the same GPUs, so drafting has capacity to use.
sed -e 's|<PREFILL_NODE>|node-a|' -e 's|<DECODE_NODE>|node-b|' \
-e 's|<PREFILL_MODEL_DIR>|/mnt/local-nvme/models|' \
-e 's|<DECODE_MODEL_DIR>|/mnt/array/models|' \
examples/recipes/kimi-k3-optimized/disaggregated-dspark/deploy.yaml | kubectl apply -f -
The placeholders are the same four as Disaggregated above, with one addition to the
requirement: Kimi-K3-DSpark/ must sit alongside Kimi-K3/ in both
directories, because both roles load the draft.
```{admonition} --speculative-config goes on BOTH roles, and that is not redundancy
- class:
warning The prefiller never samples, so a draft there looks like dead weight. That configuration was tried and it fails two independent ways:
Layer lists disagree. vLLM continues the target’s layer numbering into the
draft, so a speculating decoder registers 98 layer names and sends them all to a
prefiller that has only the target’s — KeyError: 'model.layers.93.self_attn'.
Block counts disagree. Fixing the layer lists lands straight on
pulling kv_caches ... failed: P num blocks less than D. Speculation recomputes
max_num_scheduled_tokens to reserve draft slots, so the decoder’s per-request
block accounting differs from a prefiller that does not know speculation is
happening. No layer filtering fixes that — both sides must compute it the same way.
Loading the draft on the prefiller is the price of that agreement, not an oversight. It is never run there, and both nodes need the draft on disk.
Both failures hang rather than error: all pods Ready, health checks green, restarts 0, no inference logged, and the client waits until its own timeout.
2. Prerequisites#
kubectl get nodes -o custom-columns=NODE:.metadata.name,GPU:.status.allocatable.'amd\.com/gpu'
# every manifest hardcodes `namespace: infera`, and nothing else creates it
kubectl create namespace infera --dry-run=client -o yaml | kubectl apply -f -
# on k3s, helm needs KUBECONFIG spelled out — kubectl finds it implicitly, helm does not
export KUBECONFIG=/etc/rancher/k3s/k3s.yaml
helm upgrade --install infera-operator oci://docker.io/rocm/infera-operator --version 0.1.0 \
-n infera-system --create-namespace
hf download moonshotai/Kimi-K3 --local-dir <MODEL_DIR>/Kimi-K3
hf download Inferact/Kimi-K3-DSpark --local-dir <MODEL_DIR>/Kimi-K3-DSpark # DSpark only
<MODEL_DIR> must be local NVMe. Kimi-K3’s 96 shards load in ~8 min from local
disk and ~95 min from NFS — and the slow path does not merely run late, it exceeds
the ready timeout, so the worker restarts mid-load and never finishes.
For the PD combinations each node needs its own copy, and the paths need not
match. On the fleet this was validated on, the tempting common mount (/mnt/shared)
was an NFS export of the other node’s array; df -hT <dir> on each node and take
the local device, not the convenient common name.
First start is 10–14 min: the image rebuilds its AITER JIT modules in-container on top of weight load and CUDA-graph capture.
3. Smoke test#
kubectl -n infera port-forward svc/<name>-server 18000:8000 & PF=$!
sleep 3; kill -0 $PF 2>/dev/null || { echo "port-forward failed — try another local port"; exit 1; }
# --max-time is not optional: this system's failure mode is a HANG, not an error.
curl -s --max-time 300 localhost:18000/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"kimi-k3","messages":[{"role":"user","content":"What is the capital of France?"}],
"max_tokens":1024}' | jq -r '.choices[0].message.content'
max_tokens: 1024, not 200: the model emits a variable-length reasoning preamble,
and at 200 roughly one request in four returns finish_reason: "length" having
never reached the answer — which reads like a broken deployment on a healthy one.
Keep max_tokens generous — this model spends 70+ tokens on that sentence.
On the DSpark manifests, confirm speculation engaged rather than inferring it from throughput later:
kubectl -n infera logs <worker-or-decode-pod> -c main | grep -c 'running the draft eagerly' # must be 0
On the PD manifests, a correct answer proves nothing on its own — a handoff that
fails open just re-prefills locally and returns the same text, faster. Check the
decode side: External prefix cache hit rate near 100% with Avg prompt throughput near zero.
4. Settings that are not optional#
Setting |
Why |
|---|---|
the |
selects the optimized kernels. Without them the MoE asks aiter for a kernel that was never generated — |
|
must be |
|
the ROCm counterpart of the upstream quick-start’s |
|
the draft’s weights land after the KV budget is computed; |
|
the 1800 s default is impossible on slow storage, and the worker then restarts mid-load forever — which reads as a crash loop, not as slow storage |
PD: each node needs its own local copy |
the |
PD: never point a misbehaving client at it |
the engine validates after prefill, so a rejected request has already had its KV computed and queued. 84 requests rejected for one bad field left 424 aborted Mooncake transfers and stalled valid traffic for ~20 minutes with |
ibv_devices is not installed in this image; reading its “not found” as “no
RDMA devices” produced two false TCP-fallback diagnoses here. The
ibv_get_device_list check is in the Disaggregated tab above.
5. Validation status#
Every combination was run end-to-end on the image it pins, with placeholders substituted — the files are templates.
Do not carry performance expectations across a base-image bump. Between the two images of this model the same manifest moved substantially, and in opposite directions at different concurrencies, so no single correction factor exists.
What |
Status |
|---|---|
|
validated |
|
validated; 0 restarts and the concurrency assertion never appears |
|
validated cross-node; handoff confirmed on the decode side at every concurrency exercised |
|
validated cross-node; handoff confirmed on the decode side; needs |
kvd combinations |
not built for this image |
fp8 KV cache |
not measured |