Supported features and limitations

Supported features and limitations#

2026-09-02

2 min read time

Applies to Linux

The tables list MONAI capabilities supported in the AMD ROCm 26.08 release of amd-monai 1.6.0.

Supported features#

Feature category

Feature

Notes

Inference

Sliding-window inference through SlidingWindowInferer

ROCm-optimized dynamic graph stabilization prevents graph recompilation across windows when torch.compile is active.

Inference

Patch-based and dense inference

Supported with GPU acceleration.

Network architecture

SwinUNETR (3D)

ROCm-optimized fused scaled dot-product attention auto-enabled in WindowAttention when HIP is detected.

Network architecture

DynUNet

AMD extension: use_gemm_transpose parameter and enable_gemm_transpose() method for GEMM-based transposed convolution on MI300X.

Network architecture

VISTA3D, SegResNet, UNETR, BasicUNet, and others

Supported without ROCm-specific modifications.

Transforms

GPU-accelerated transforms

Supported for spatial, intensity, and elastic transforms through the PyTorch HIP backend.

Data loading

NIfTI, DICOM, MHA, MHD, PNG, and JPEG

CPU-based I/O with GPU transfer through DataLoader.

Data loading

Whole-slide image reading through WSIReader

GPU-accelerated with amd-hipcim backend when backend="cuCIM".

GPU acceleration

Mixed precision (BF16 and FP16)

BF16 is preferred on MI300X and MI355X through PyTorch AMP (torch.amp.autocast).

GPU acceleration

torch.compile graph optimization

Supported. First-call compilation latency is expected.

Model Zoo

MONAI Bundle format

Supported. AMD overlay mechanism adds ROCm optimizations at runtime without modifying upstream bundles.

Model Zoo

Five validated bundles

VISTA3D, SwinUNETR BTCV, Whole Body CT, Spleen DeepEdit, and Pancreas DiNTS.

Metrics and losses

DiceLoss, DiceCELoss, FocalLoss, Hausdorff distance, MeanIoU

Supported.

Interoperability

NumPy, ITK, SimpleITK array conversions

Supported through the CPU bridge.

Interoperability

amd-cupy

Supported with amd-cupy 14.1.1 or later.

Foundation model

EXAONEPath 2.0

Computational pathology foundation model validated on AMD hardware. ViT-based WSI patch inference through Hugging Face.

Limitations#

Limitation

Details

GPU direct storage through KvikIO or cuFile

Not supported on ROCm. Standard CPU-mediated I/O is used instead.

rocTX and NVTX profiling markers

MONAI NVTX-based profiling annotations are not functional on ROCm. Use rocprof or Omniperf directly.

CuPy version

Requires amd-cupy 14.1.1 or later. Standard NVIDIA CuPy packages are not compatible on ROCm.

hipCIM version

WSI support requires amd-hipcim 26.06.00 or later.

torch.compile first-call latency

Expect 30 to 120 seconds of compilation on the first call when torch.compile is enabled. Subsequent calls use the cached graph.

Network architecture support matrix#

Architecture

ROCm support

AMD extensions

Tasks

SwinUNETR

Supported

Fused SDPA auto-enable

CT and MRI segmentation, whole-body

DynUNet

Supported

GEMM-based transpose conv

Segmentation, nnU-Net backbone

VISTA3D

Supported

BF16 and torch.compile overlay

Universal volumetric segmentation

SegResNet

Supported

None

Brain tumor segmentation (BraTS)

UNETR

Supported

None

Transformer-based segmentation

BasicUNet

Supported

None

General encoder-decoder

DiNTS (NAS)

Supported

None

NAS-discovered segmentation

EXAONEPath 2.0

Supported through Hugging Face

None

Computational pathology, WSI