Kubernetes deployment#
Infera’s production path is Kubernetes-native: install the platform (operator +
NATS + LeaderWorkerSet controller) with Helm, submit an InferaDeployment
CRD, and let the operator reconcile it into Deployments / LeaderWorkerSets,
Services, and the KV-event plane. Ready-to-fill InferaDeployment templates
(single-node, prefill/decode, multi-node TP, GAIE) live under
examples/k8s-deployments/.
Prerequisites#
A Kubernetes cluster (v1.24+) with ≥ 1 ROCm node, and
kubectl+helm.The AMD GPU device plugin, so GPUs are schedulable as
amd.com/gpu. If it’s not installed yet, follow the AMD GPU Operator setup guide or the device plugin README.kubectl describe node | grep amd.com/gpu # Capacity should be > 0
This is the hard prerequisite. etcd is not needed — the operator’s default discovery is the in-cluster API.
Your engine image (
infera/engine-{sglang,vllm,atom}) pullable by the nodes; a per-namespaceimagePullSecretfor a private registry.Model data reachable inside the pod (a
hostPath/PVC viaextraPodSpec, baked into the image, or an HF download with a token secret).
1. Install the platform#
One script installs the cluster dependencies, each as an idempotent
helm upgrade --install: the LeaderWorkerSet controller (multi-node workers),
NATS / JetStream (KV-event + request transport), and the Infera operator
(with its CRD + RBAC). Clone the Infera repo first if you haven’t already:
git clone https://github.com/AMD-AGI/Infera.git && cd Infera
deploy/scripts/deploy-k8s.sh # LWS + NATS + operator (namespace: infera-system)
# toggles: INSTALL_NATS=false · INSTALL_GATEWAY=true (kgateway/GAIE) · OPERATOR_IMAGE_TAG=…
# deploy/scripts/deploy-k8s.sh --help for the full list
Verify the operator and CRD are up:
kubectl get crd inferadeployments.infera.amd.com
kubectl -n infera-system get deploy,pod # operator (+ NATS) Running
2. Deployment resources#
Resource |
What it is |
|---|---|
|
the canonical CR — one inference graph: a |
Deployment / LeaderWorkerSet |
the per-pool workloads the operator creates (LWS when |
Service + NATS |
a ClusterIP for the server and the KV-event broker — created for you. |
The full field reference is on the Operator page.
3. Deploy your first model#
Apply an InferaDeployment. Aggregated (mixed) serving first:
# infera-qwen.yaml
apiVersion: infera.amd.com/v1alpha1
kind: InferaDeployment
metadata:
name: qwen
spec:
backendFramework: sglang
image: rocm/infera:sglang-v0.1.1
nats: {deploy: true}
services:
server:
componentType: server
replicas: 1
args: ["--router-tokenizer-path", "Qwen/Qwen3-0.6B"]
worker:
componentType: worker
role: mixed
replicas: 2
resources: {gpu: 1, gpuType: amd.com/gpu}
args: ["--model-path", "Qwen/Qwen3-0.6B", "--tp-size", "1"]
kubectl apply -f infera-qwen.yaml
For PD (prefill/decode pools) and multi-node (numberOfNodes > 1 →
LeaderWorkerSet), see the full disaggregated CRD example on the
Operator page.
4. Monitor#
kubectl get idep qwen -w # BACKEND / STATE → ready
kubectl get pods -l infera.amd.com/deployment=qwen # server + workers Running
kubectl logs -f deploy/qwen-worker # engine model-load progress
5. Send a request#
kubectl port-forward svc/qwen-server 8000:80 &
curl localhost:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"Qwen/Qwen3-0.6B",
"messages":[{"role":"user","content":"1+1=?"}],
"max_tokens":50}'
Scaling & upgrades#
Scale —
kubectl scale(or editreplicasin the CR) on a worker pool; the operator reconciles the Deployment/LWS.Upgrade — change
image/argsin the CR and re-apply; native Deployment/LWS rollout drains in-flight requests.Logs / registry —
kubectl logs -l infera.amd.com/deployment=<name>.
Deployment templates#
Instead of hand-writing the CR above, start from the ready-to-fill templates in
examples/k8s-deployments/ — one per
topology (single-node aggregated, prefill/decode disaggregated, multi-node TP,
and GAIE), each with a shared model-store PVC and documented <PLACEHOLDERS>.
Gateway API (GAIE)#
To front the fleet with a Kubernetes Inference Gateway — an Endpoint Picker
(EPP) + InferencePool + HTTPRoute + a per-worker frontend sidecar — set
spec.gaie on the CR and install the gateway stack:
INSTALL_GATEWAY=true deploy/scripts/deploy-k8s.sh
The server then runs in --router-mode direct, honoring the EPP’s per-request
worker pick. See Routing & transport.