Run a Hyperloom optimization#
This topic assumes you have already completed installation. If you haven’t, go through Hyperloom Quickstart then return here to launch your first run.
Launch from Cursor#
Open the Hyperloom workspace in Cursor, then paste the following prompt into Cursor Chat, filling in your workload details:
Note
The prompt includes install.sh. This is intentional: Cursor runs in its own
shell process, which does not inherit the environment you sourced during
installation. The agent must re-source the env files and re-run install.sh in
its own context before launching the optimizer. Because install.sh is
idempotent, the second run is fast and safe.
@src/hyperloom/inference_optimizer/SKILL.md
Optimize inference for this workload:
- Model: /path/to/your/model
- Framework: sglang
- GPU: MI300X
- TP: 1
- CONC: 64
- ISL: 1024
- OSL: 1024
- Goal: improve throughput by at least 10%
- Budget: 24 hours
Before launch, run exactly:
bash "$REPO_ROOT/src/hyperloom/inference_optimizer/assets/install.sh"
source '/path/to/hyperloom-run/runtime/kernel-agent.env.sh'
export USER_DATA_PATH='/path/to/hyperloom-run'
Requirements:
1. Report the session ID, log path, PID, and initial health check result.
2. Monitor the process every 300s until the optimization is complete or failed.
Field |
Meaning |
How to choose |
|---|---|---|
|
Tensor-parallel size — number of GPUs the model is sharded across |
Must match the number of GPUs in your server node (for example, |
|
Concurrent requests — baseline benchmark concurrency ( |
Set to your target concurrency; the post-run concurrency sweep separately measures a ladder (default |
|
Input sequence length — tokens in each request’s prompt |
Match your production workload; |
|
Output sequence length — tokens generated per response |
Match your production workload; |
See src/hyperloom/inference_optimizer/SKILL.md
for the full prompt field reference (every field maps to a CLI flag defined in
cli/parser.py).
Monitor the run#
The agent reports a session ID, log path, and PID, then polls until the run
completes. Under the hood it walks the phase chain
PRELUDE → FRAMEWORK_AGENT → EXPLORE → KERNEL_AGENT → SWEEP → CLOSE; see
Hyperloom optimization loop for what
happens in each phase.
Resume an interrupted session#
Paste this prompt into Cursor Chat to resume an existing session:
@src/hyperloom/inference_optimizer/SKILL.md
Resume the existing Hyperloom optimization session.
Requirements:
1. Launch `python -m hyperloom.inference_optimizer.cli optimize --resume`; do not start a new session.
2. Do not pass `--model`; read the model and workload from the saved manifest.
3. Before launching, verify `manifest.json` and `state.json` exist.
4. Report the log path, PID, health check, current phase, cumulative gain, and best config.
5. Monitor the process every 300s until the optimization is complete or failed.
Output and artifacts#
When the loop exits, Hyperloom writes the final report, reproducible session
artifacts, and session_breakdown.json to your session directory. The three
fields to read first are:
Field |
What it tells you |
|---|---|
|
Validated end-of-session serving throughput — the headline number for SGLang / vLLM / Atom |
|
Validated gain over baseline |
|
Ordered list of changes that make up the final optimized stack |
For scriptable diffusion workloads (--framework xdit), the headline metric is
final.e2el_mean_ms (lower is better) and reports use img/s as the throughput
unit.
For the full schema — useful if you are building a dashboard, reporting
pipeline, or downstream integration on top of this file — see
session_breakdown.json integration in Hyperloom.
Troubleshooting#
If a run fails on first launch, see Troubleshooting Hyperloom.