Run xDiT diffusion inference tests with CVS#
2026-09-24
6 min read time
CVS provides five xDiT suites under cvs/tests/inference/xdit/. Each suite is a
separate pytest module; pick the one that matches your topology and launcher, then point
--config_file at a template from cvs/input/config_file/inference/xdit/.
Single-node suites run one independent docker+torchrun job on every node in the cluster file (full model on each node).
Distributed suites run one coordinated torchrun job across
nnodes(nnodes >= 2), usingserver_node_listwhen set.
FLUX.1-dev and FLUX.2-dev share the pytorch_xdit_flux_dev_* suites; choose the matching
flux1 or flux2 JSON. Models must already be staged on every participating node
(no runtime Hugging Face downloads).
Config reference: xDiT inference benchmark configuration for Cluster Validation Suite (CVS).
Test suites#
The following suites are available.
CVS suite name |
Source module |
What it runs |
|---|---|---|
|
|
FLUX.1 ( |
|
|
Unified FLUX.1 / FLUX.2 torchrun across |
|
|
WAN 2.2 I2V native ( |
|
|
WAN Diffusers xFuser ( |
|
|
Unified WAN Diffusers xFuser torchrun across |
Set up config#
Follow these steps to set up the xDiT configuration.
List available xDiT templates:
cvs config list inference/xdit
Copy the configuration for your workload. FLUX.1-dev and FLUX.2-dev share the same suites; copy the matching
flux1orflux2JSON:cvs config copy inference/xdit/mi3xx_pytorch_xdit_flux1_dev_single.json \ --output ~/cvs_workspace/inference/xdit/mi3xx_pytorch_xdit_flux1_dev_single.json cvs config copy inference/xdit/mi3xx_pytorch_xdit_wan22_14b_single.json \ --output ~/cvs_workspace/inference/xdit/mi3xx_pytorch_xdit_wan22_14b_single.json
Copy a cluster file (GPU compute nodes only):
cvs config copy cluster_container.json --output ~/cvs_workspace/cluster.json
Edit the config — set
container_image,model_repo/hf_home,nnodesfor distributed runs, and replace every<changeme>. Resolve{user-id}/{home}or leave them for CVS to expand.
Shipped config templates:
Config file |
Use with suite |
|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Note
FLUX.2 configs bind-mount cvs/lib/inference/xdit/scripts/flux2_example.py when the
image does not ship /app/external/xdit/examples/flux2_example.py. WAN Diffusers
suites require model_repo as an absolute host path on every node and typically
mount cvs/lib/inference/xdit/scripts/wan_i2v_example.py.
On shared clusters, skip aggressive docker prune during cleanup:
export CVS_PYTORCH_XDIT_SKIP_DOCKER_SYSTEM_PRUNE=1
Run tests#
List stages in a suite:
cvs list pytorch_xdit_flux_dev_single
pytorch_xdit_flux_dev_single stages#
Available tests in pytorch_xdit_flux_dev_single:
- test_cleanup_stale_containers
- test_verify_hf_cache_or_download
- test_run_flux1_benchmark
- test_parse_and_validate_results
Example run (FLUX.1-dev; use the flux2 JSON for FLUX.2-dev):
cvs run pytorch_xdit_flux_dev_single \
--cluster_file ~/cvs_workspace/cluster.json \
--config_file ~/cvs_workspace/inference/xdit/mi3xx_pytorch_xdit_flux1_dev_single.json \
--html ~/cvs_results/pytorch_xdit_flux1_single.html --self-contained-html \
--log-file /tmp/pytorch_xdit_flux1_single.log -vvv
pytorch_xdit_flux_dev_distributed stages#
Available tests in pytorch_xdit_flux_dev_distributed:
- test_cleanup_stale_containers
- test_verify_hf_cache_or_download
- test_verify_parallelism_config
- test_run_flux1_benchmark
- test_parse_and_validate_results
test_verify_parallelism_config checks that
ulysses × ring × pipefusion × tp × dp == nnodes × torchrun_nproc.
Example run:
cvs run pytorch_xdit_flux_dev_distributed \
--cluster_file ~/cvs_workspace/cluster.json \
--config_file ~/cvs_workspace/inference/xdit/mi3xx_pytorch_xdit_flux1_dev_distributed.json \
--html ~/cvs_results/pytorch_xdit_flux1_distributed.html --self-contained-html \
--log-file /tmp/pytorch_xdit_flux1_distributed.log -vvv
pytorch_xdit_wan22_14b_single stages#
Available tests in pytorch_xdit_wan22_14b_single:
- test_cleanup_stale_containers
- test_verify_hf_cache_or_download
- test_run_wan22_benchmark
- test_parse_and_validate_results
Example run:
cvs run pytorch_xdit_wan22_14b_single \
--cluster_file ~/cvs_workspace/cluster.json \
--config_file ~/cvs_workspace/inference/xdit/mi3xx_pytorch_xdit_wan22_14b_single.json \
--html ~/cvs_results/pytorch_xdit_wan22.html --self-contained-html \
--log-file /tmp/pytorch_xdit_wan22.log -vvv
pytorch_xdit_wan22_14b_diffusers_single stages#
Available tests in pytorch_xdit_wan22_14b_diffusers_single:
- test_cleanup_stale_containers
- test_verify_model_on_nodes
- test_run_wan22_diffusers_benchmark
- test_parse_and_validate_results
Example run:
cvs run pytorch_xdit_wan22_14b_diffusers_single \
--cluster_file ~/cvs_workspace/cluster.json \
--config_file ~/cvs_workspace/inference/xdit/mi3xx_pytorch_xdit_wan22_14b_diffusers_single.json \
--html ~/cvs_results/pytorch_xdit_wan22_diffusers_single.html --self-contained-html \
--log-file /tmp/pytorch_xdit_wan22_diffusers_single.log -vvv
pytorch_xdit_wan22_14b_diffusers_distributed stages#
Available tests in pytorch_xdit_wan22_14b_diffusers_distributed:
- test_cleanup_stale_containers
- test_verify_model_on_nodes
- test_verify_parallelism_config
- test_run_wan22_diffusers_benchmark
- test_parse_and_validate_results
test_verify_parallelism_config checks that ulysses_size × ring_size == nnodes × torchrun_nproc.
Example run:
cvs run pytorch_xdit_wan22_14b_diffusers_distributed \
--cluster_file ~/cvs_workspace/cluster.json \
--config_file ~/cvs_workspace/inference/xdit/mi3xx_pytorch_xdit_wan22_14b_diffusers_distributed.json \
--html ~/cvs_results/pytorch_xdit_wan22_diffusers_distributed.html --self-contained-html \
--log-file /tmp/pytorch_xdit_wan22_diffusers_distributed.log -vvv
Direct pytest invocation#
Each module can also be run with pytest:
pytest cvs/tests/inference/xdit/pytorch_xdit_flux_dev_single.py \
--cluster_file ~/cvs_workspace/cluster.json \
--config_file ~/cvs_workspace/inference/xdit/mi3xx_pytorch_xdit_flux1_dev_single.json \
--html ~/cvs_results/pytorch_xdit_flux1_single.html
Read the results#
With --html, CVS writes a pytest HTML report. Benchmark pass/fail uses the docker
exit code plus parsed artifacts and GPU-specific thresholds (mi300x, mi350,
mi355, or auto).
Key stages to watch:
Cleanup —
test_cleanup_stale_containersstops the named container (and{container_name}-rankNon distributed suites). It also runsdocker system pruneunlessCVS_PYTORCH_XDIT_SKIP_DOCKER_SYSTEM_PRUNE=1.Model preflight —
test_verify_hf_cache_or_download(FLUX and WAN native) ortest_verify_model_on_nodes(WAN Diffusers). Fails if the model is missing on any participating node.Parallelism —
test_verify_parallelism_config(distributed suites only).Benchmark —
test_run_flux1_benchmark,test_run_wan22_benchmark, ortest_run_wan22_diffusers_benchmark.Parse —
test_parse_and_validate_resultscompares average latency toexpected_results.
Family |
Metric |
Threshold key |
Artifacts |
|---|---|---|---|
FLUX |
average |
|
|
WAN native |
average |
|
step JSONs, |
WAN Diffusers |
average epoch / pipe time from |
|
|
Single-node output dirs use the cluster SSH target, for example
${output_base_dir}/flux_<target>_outputs or wan_22_<target>_outputs.
Distributed runs write to the rank-0 target directory.