Magpie#
Magpie is a lightweight, general-purpose framework for evaluating GPU kernel correctness and performance on AMD (HIP) GPUs. It exposes three evaluation modes — Analyze, Compare, and Benchmark — plus framework-level (vLLM / SGLang / Atom) benchmarking with built-in TraceLens trace analysis.
Within Hyperloom, Magpie is the benchmark engine. The kernel agent and the
optimization loop drive Magpie to spin up a serving framework, run the workload,
collect traces, and emit a structured benchmark_report.json file; those traces are
the input that TraceLens then analyzes. Magpie relies on
IntelliKit for some low-level GPU profiling tools.
Source: AMD-AGI/Magpie
License: MIT
Role in Hyperloom#
The optimization loop drives Magpie’s benchmark mode as a subprocess. The
command line is built by build_benchmark_command() /
MagpieBackend.build_command() in
src/hyperloom/orchestrator/actions/executors/benchmark_backend.py; callers
such as _grid_runner.py and baseline.py invoke it to launch one run per
variant:
cmd = [
magpie_python, "-m", "Magpie", "-v", "benchmark",
"--benchmark-config", str(config_path),
"--output-dir", str(output_dir),
"--run-mode", "local",
]
Each run produces a benchmark_report.json that Hyperloom parses to extract
throughput/measurements and pick winners. To make concurrent benchmark runs
robust, src/hyperloom/orchestrator/actions/executors/_magpie_patcher.py
applies an idempotent, atomic-write patch to Magpie’s installed benchmarker.py
(_prepare_benchmark_scripts) so a concurrent reader never sees a half-copied
script. See Hyperloom optimization loop for more information.
Magpie documentation#
For detailed documentation on Magpie, see Magpie on ROCm Docs.