Compiler and pipeline#
FlyDSL includes a JIT compiler that traces Python kernel functions into MLIR and lowers them through the Fly dialect pipeline to GPU binaries.
@flyc.kernel and @flyc.jit#
The primary API for defining and compiling kernels:
import flydsl.compiler as flyc
import flydsl.expr as fx
@flyc.kernel
def my_kernel(A: fx.Tensor, B: fx.Tensor, n: fx.Constexpr[int]):
tid = fx.thread_idx.x
bid = fx.block_idx.x
# ... kernel body using layout ops ...
@flyc.jit
def launch(A: fx.Tensor, B: fx.Tensor, n: fx.Constexpr[int],
stream: fx.Stream = fx.Stream(None)):
grid_x = (n + 255) // 256
my_kernel(A, B, n).launch(
grid=(grid_x, 1, 1),
block=(256, 1, 1),
stream=stream,
)
@flyc.kernelcompiles the function body into agpu.funcinside agpu.module. It uses AST rewriting to trace Python code into MLIR IR.@flyc.jitwraps a host-side function that constructs and launches kernels. On first call it triggers JIT compilation; subsequent calls with the same type signature use a cached compiled artifact.
Compilation flow#
On first call, @flyc.jit runs the following pipeline:
AST rewriting: The Python source is parsed and rewritten to emit MLIR ops.
MLIR module construction: The kernel body is traced into
fly,gpu,arith,scf,memref, andvectordialect ops.Fly pass pipeline: The module is lowered through three pass stages, defined in
RocmBackend._pipeline_parts()(python/flydsl/compiler/backends/rocm.py). See Architecture & compilation pipeline guide §3 for the per-pass table.pre_binary_fragments(Fly → ROCDL):fly-rewrite-func-signaturefly-canonicalizefly-layout-loweringfly-int-swizzle-simplifycanonicalizefly-convert-atom-call-to-ssa-formfly-promote-regmem-to-vectorssaconvert-fly-to-rocdlcanonicalizegpu.module(convert-scf-to-cf, cse, convert-rocdl-fastmath-ops, convert-gpu-to-rocdl{chipset=gfxNNN ...}, fly-rocdl-cluster-attr)
binary_prep_fragments(→ LLVM):rocdl-attach-target{chip=gfxNNN ...}convert-scf-to-cfconvert-cf-to-llvmgpu-to-llvm{use-bare-pointers-...=true}convert-vector-to-llvmconvert-arith-to-llvmconvert-func-to-llvmreconcile-unrealized-castsensure-debug-info-scope-on-llvm-func(optional, gated byFLYDSL_DEBUG_ENABLE_DEBUG_INFO)
binary_fragment:gpu-module-to-binary{format=fatbin opts="..."}
Cached artifact: The compiled binary is cached to disk (
~/.flydsl/cache/) keyed by the compiler toolchain hash and kernel type signature.
Ahead-of-time C export#
CompiledFunction.export_to_c(file_path, file_name, function_prefix="")
writes a target-specific object file and a C/C++ header. The object contains
the GPU binary and a small backend adapter, so the deployed executable or
shared library does not need FlyDSL or Python. The generated header records the
backend and system link flags; for ROCm these are
-lamdhip64 -pthread -ldl. The linker must also be able to find the HIP SDK
library. For a nonstandard ROCm installation,
add its library directory, for example:
cc -shared -o libkernel.so kernel.o -L/path/to/rocm/lib -lamdhip64 -pthread -ldl
The metadata records libamdhip64.so as a link-time name. The versioned
runtime dependency (the ELF DT_NEEDED name) is determined by the library
selected when the final executable or shared library is linked. The deployed
system’s dynamic loader must be able to find that versioned HIP library.
The header exposes two calling styles:
<symbol>__module_init,<symbol>__module_loadand<symbol>__module_unloadprovide explicit lifecycle and device control. After loading,<symbol>_callprovides a typed wrapper around the packed entry point.<symbol>_call_autois the simple path. It idempotently initializes the module and loads it on the current device before each call. It intentionally does not unload automatically, so asynchronous launches remain safe.
The current exporter targets 64-bit little-endian Linux ELF hosts and uses
GNU-compatible relocatable linking and objcopy tools. The ROCm adapter
requires HIP and pthread. Exported objects are tied to the host ABI and GPU
target used during compilation. They can be moved and linked independently,
but must run on a compatible backend and GPU architecture.
Tensor arguments#
Use flyc.from_dlpack to convert PyTorch tensors into FlyDSL tensor
descriptors with layout metadata:
import torch
import flydsl.compiler as flyc
A = torch.randn(1024, device="cuda", dtype=torch.float32)
B = torch.empty_like(A)
tA = flyc.from_dlpack(A).mark_layout_dynamic(
leading_dim=0, divisibility=4
)
tB = flyc.from_dlpack(B)
launch(tA, tB, A.numel(), stream=torch.cuda.Stream())
Precompilation and C export#
flyc.compile(launcher, *specialization_args, **specialization_kwargs)
compiles a @flyc.jit launcher for one argument signature and returns a
CompiledFunction. The result is callable with new runtime values of the
same signature, and Constexpr values remain fixed to the specialization
used at compile time.
The same compiled specialization can be exported as a position-independent host object with a generated C header:
from pathlib import Path
output = Path("build/aot")
output.mkdir(parents=True, exist_ok=True)
compiled = flyc.compile(launch, tA, tB, A.numel())
compiled.export_to_c(
file_path=output,
file_name="vector_add",
function_prefix="flydsl_vector_add",
)
export_to_c writes vector_add.o and vector_add.h into the existing
output directory. The object embeds the FlyDSL backend adapter; link the final
binary against the backend system libraries listed in the header. The header
declares the packed entry point, typed inline call helpers, module
initialization/load/unload functions, and embedded ABI metadata. When
function_prefix is omitted, file_name is also used as the exported C
symbol.
Passing live device arguments to flyc.compile preserves its normal initial
launch. For an offline or CPU-only build host, use null pointer wrappers created
with flyc.from_c_void_p (or compatible non-device tensor placeholders) and
set the target architecture explicitly, for example with ARCH=gfx950.
Export currently supports tensor, pointer, scalar, structure, and stream
arguments whose lowered types have a supported C ABI. Unsupported launchers
report the export error when export_to_c is called.
ROCDL operations#
The flydsl.expr.rocdl module provides AMD-specific operations:
The universal exported entries follow API Stability. The gfx1250
WMMAScale and TDM helpers below are target-specific implementation APIs and
remain unstable.
fx.rocdl.make_buffer_tensor – create buffer resource descriptor from tensor (CDNA buffer copy)
fx.rocdl.BufferCopy32b / BufferCopy128b – buffer copy atoms
fx.rocdl.MFMA – MFMA instruction atoms (CDNA3/CDNA4; for example,
MFMA(16, 16, 4, fx.Float32))fx.rocdl.WMMA / fx.rocdl.WMMAScale – wave32 WMMA MMA atoms;
WMMAis arch-dispatched (gfx11 / gfx120x RDNA4 / gfx1250),WMMAScaleis the gfx1250 E8M0 MX-scaled formfx.rocdl.make_tdm_atom / fx.rocdl.TDM – gfx1250 TDM async Global↔LDS whole-tile copy atom (1–5D; base from the copy operand, per-dim extent/stride/imm_offset/mask as atom state)
fly-opt CLI#
The fly-opt tool is a command-line interface for running MLIR passes on
.mlir files:
fly-opt --fly-canonicalize input.mlir
fly-opt --fly-layout-lowering input.mlir
fly-opt --help