xDiT inference benchmark configuration for Cluster Validation Suite (CVS)#
2026-09-24
7 min read time
CVS ships five xDiT suites under cvs/tests/inference/xdit/. Each suite reads a JSON
file from cvs/input/config_file/inference/xdit/. There is no separate threshold file;
expected_results live inside benchmark_params.
Single-node templates run one independent docker+torchrun job on every node in the cluster file.
Distributed templates run one coordinated torchrun job (
nnodes >= 2). Replace every<changeme>(NCCL/network fields) before running.
See Run xDiT diffusion inference tests with CVS for more information on running these tests.
Note
{user-id}and{home}in path strings are resolved at runtime.Models must already be staged on every participating node. Prefer an absolute path in
model_repo; a Hugging Face repo id requires a pre-populated cache underhf_home.FLUX.1-dev and FLUX.2-dev share
pytorch_xdit_flux_dev_*; pick the matching JSON.
Configuration files#
Each xDiT suite maps to one or more JSON templates; select the file that matches your model and execution mode.
Config file |
Use with suite |
|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Copy a template:
cvs config list inference/xdit
cvs config copy inference/xdit/mi3xx_pytorch_xdit_flux1_dev_single.json \
--output ~/cvs_workspace/inference/xdit/mi3xx_pytorch_xdit_flux1_dev_single.json
File structure#
Every template has two top-level keys:
Key |
Description |
|---|---|
|
Image, model path, output dir, optional distributed rendezvous/NCCL, |
|
|
Example: FLUX.1-dev single-node#
mi3xx_pytorch_xdit_flux1_dev_single.json (abbreviated)
{
"config": {
"container_image": "amdsiloai/pytorch-xdit:v25.11.2",
"container_name": "flux-benchmark",
"hf_token_file": "{home}/.hf_token",
"hf_home": "{home}/.cache/huggingface",
"output_base_dir": "{home}/cvs_flux_output",
"model_repo": "black-forest-labs/FLUX.1-dev",
"model_rev": "",
"container_config": {
"device_list": ["/dev/dri", "/dev/kfd"],
"volume_dict": {},
"env_dict": {}
}
},
"benchmark_params": {
"flux1_dev_t2i": {
"prompt": "A small cat",
"num_inference_steps": 25,
"num_repetitions": 25,
"height": 1024,
"width": 1024,
"ulysses_degree": 8,
"ring_degree": 1,
"use_torch_compile": true,
"torchrun_nproc": 8,
"expected_results": {
"auto": { "max_avg_pipe_time_s": 10.0 },
"mi300x": { "max_avg_pipe_time_s": 3.0 },
"mi350": { "max_avg_pipe_time_s": 2.0 },
"mi355": { "max_avg_pipe_time_s": 7.0 }
}
}
}
}
General config parameters#
The following parameters appear in the config block and are common across all xDiT templates.
Parameter |
Example |
Description |
|---|---|---|
|
|
Image with PyTorch xDiT. FLUX.2 and WAN Diffusers templates use |
|
|
Docker name (distributed ranks use |
|
|
Hugging Face token for gated models (FLUX.2 chat template). |
|
|
Host HF cache (mounted at |
|
|
Host directory for |
|
HF id or |
Repo id (offline cache) or absolute host path (preferred). WAN Diffusers requires an absolute path. |
|
snapshot hash or |
HF snapshot id when using cache mode. WAN native template pins |
|
|
GPU device nodes passed into the container. |
|
host → container map |
Extra bind mounts. FLUX.2 mounts |
|
|
Extra environment variables inside the container. |
Distributed config fields#
Present on *_distributed.json templates (FLUX.1, FLUX.2, WAN Diffusers):
Parameter |
Description |
|---|---|
|
Participating node count (must be |
|
torchrun rendezvous (port default |
|
NCCL InfiniBand/RoCE devices and GID index. |
|
Ethernet interfaces for socket/Gloo fallback. |
|
NCCL log level (templates use |
benchmark_params.flux1_dev_t2i#
Used by all four FLUX templates. FLUX.2 sets model_type: flux2.
Parameter |
Example |
Description |
|---|---|---|
|
|
FLUX.2 only. Selects |
|
|
Generation prompt and RNG seed. |
|
|
FLUX.2 guidance (FLUX.1 templates omit this). |
|
|
Denoising steps (FLUX.1 |
|
|
Text encoder sequence length. |
|
|
Disable resolution binning. |
|
|
Warmup then measured repetitions. |
|
|
Output image size. |
|
|
Sequence-parallel layout. Product with pipefusion/TP/DP must equal |
|
|
Enable |
|
|
Processes (GPUs) per node. |
|
|
Pass/fail on average |
benchmark_params.wan22_i2v_a14b#
This block configures the WAN 2.2 image-to-video benchmark and supports both native WAN and Diffusers execution paths.
Native WAN (mi3xx_pytorch_xdit_wan22_14b_single.json)#
Runs /app/Wan2.2/run.py. Threshold metric is max_avg_total_time_s.
Parameter |
Example |
Description |
|---|---|---|
|
(long I2V prompt) |
Image-to-video prompt. |
|
|
Frame size. |
|
|
Number of video frames. |
|
|
Measured steps after compile/warmup. |
|
|
Enable compile on the native launcher. |
|
|
GPUs per node. |
|
|
Average |
Diffusers xFuser WAN#
mi3xx_pytorch_xdit_wan22_14b_diffusers_*.json additionally set:
Parameter |
Description |
|---|---|
|
|
|
|
|
In-container path to |
|
Generate an in-container input image when true. |
|
Install video encode deps inside the container when true. |
|
|
|
|
|
|
|
Fail parse if |
|
Denoising and warmup (Diffusers templates: |
|
Parallel layout. Distributed: product must equal |
|
|
Volume mounts#
FLUX.1 / WAN native templates ship an empty volume_dict. Bind-mount models via
model_repo as an absolute path, or rely on hf_home.
FLUX.2 mounts the in-tree example when the image lacks it:
{
"volume_dict": {
"/home/{user-id}/cvs/cvs/lib/inference/xdit/scripts/flux2_example.py": "/benchmark/flux2_example.py"
}
}
WAN Diffusers mounts the xFuser example:
{
"volume_dict": {
"/home/{user-id}/cvs/cvs/lib/inference/xdit/scripts/wan_i2v_example.py": "/benchmark/wan_i2v_example.py"
}
}
Adjust the host path to your CVS checkout.
Performance metrics#
GPU type is detected from rocm-smi. Lookup order: exact key → auto.
FLUX — average
pipe_timevsmax_avg_pipe_time_s; artifactsresults/timing.jsonandflux_*.png.WAN native — average
total_timevsmax_avg_total_time_s;rank0_step*.jsonandvideo.mp4.WAN Diffusers — average pipe/epoch time vs
max_avg_pipe_time_s;results/timing.jsonandresults/video_i2v.mp4.
Shipped numbers are starting points; tune expected_results for your stack before production gating.
Troubleshooting#
- ``/dev/kfd not found``
Run on GPU compute nodes, not login nodes.
- Container image not found locally
docker pullthe configuredcontainer_imageon every execution node.- Local model path not found
Stage weights on every participating node. Diffusers WAN requires
model_repoas an absolute path.- Parallel degree product != world_size
Align
ulysses/ring(and FLUX pipefusion/TP/DP) withnnodes × torchrun_nproc.- Missing ``timing.json`` / ``video.mp4``
The benchmark docker exit code was non-zero or artifacts were written elsewhere; inspect the log tail on the failing node.