tracelens-analysis-orchestrator#
The goal of this skill is to teach your AI agent to run the TraceLens agentic
analysis workflow on a PyTorch trace and produce a prioritized stakeholder
report (analysis.md) with recommendations for optimizing individual kernels,
kernel fusion opportunities, and system-level optimizations.
What you’ll end up with: a single user-facing report at
<output_dir>/analysis.md with an executive summary, ranked P-items by impact score,
and a detailed analysis section with evidence tables and resolution guidance.
Expect this process to take 5–30 minutes depending on trace size.
Analysis modes at a glance#
The tables below summarize which analysis modes and workloads TraceLens supports.
Mode |
Traces |
Best for |
|---|---|---|
Standalone |
1 |
Roofline analysis on a single trace (default) |
Comparative |
2 |
Gap analysis: primary trace vs. a faster reference trace |
Workload |
Standalone |
Comparative |
|---|---|---|
Eager (training or eager inference) |
Yes |
Yes |
Graph + capture (vLLM or SGLang) |
Yes |
No |
Graph only |
No |
No |
For vLLM or SGLang inference traces, the canonical collection guide in TraceLens Inference Analysis provides detailed instructions about collecting traces. Including capture mode graphs will produce better results but might require patching vLLM or SGLang.
Prerequisites#
Confirm that your software and data meet these requirements before you start.
Software
Node.js 22 or later, required for the
skillsCLI (npx skills ...)An agentic runner: Cursor (chat or
agentCLI) or Claude CodeModel: Claude Opus 4.7 High (
claude-opus-4-7-high). The orchestrator and all 13 sub-agents are tuned for this model.
Data
A PyTorch profiler trace (
.jsonor.json.gz) from a representative steady-state window (post-warmup). A single rank’s trace is sufficient for per-rank analysis.The platform of the first trace (for example,
MI300X,MI325X), used for roofline limits and hardware reference in the report
Step 1: Check which skills are available#
Start with a clean agent session so skill discovery reflects only what you installed.
Run
claude "which skills do you see?". You should see a list of skills that should not include anything related to TraceLens.Make sure there is no
AGENTS.mdfile on your local folder.
Step 2: Enable your agent to see tracelens-analysis-orchestrator#
Run the npx commands below in your terminal, not inside Claude.
Claude Code: install with the skills CLI:
npx skills add amd/skills --skill tracelens-analysis-orchestrator --agent claude-code
Cursor: install the AMD Skills plugin or add the skill manually:
npx skills add amd/skills --skill tracelens-analysis-orchestrator --agent cursor
Confirm the skill is visible:
claude "Which skills can you see?" --model opus
You should now see tracelens-analysis-orchestrator in the list.
Step 3: Run standalone analysis (interactive)#
This is the recommended first run: one trace, default analysis mode (training and non-vLLM or SGLang eager inference).
Cursor chat#
Open a Cursor chat with Claude Opus 4.7 High and send:
Follow the analysis orchestrator installed with TraceLens and run the full agentic
analysis workflow on <path_to_trace.json>
When prompted, provide:
Input |
Example |
|---|---|
Platform |
|
Analysis mode |
|
Environment |
local, local+venv, cluster, or cluster+container |
Node, container, or venv |
only if running remotely |
Output directory |
|
For vLLM or SGLang traces, choose analysis mode inference and specify execution
mode (eager or graph replay + capture). Graph replay + capture also requires a
capture folder path.
Claude Code#
claude --model opus
Then prompt:
Follow the analysis orchestrator and run the full agentic analysis workflow on
<path_to_trace.json> with platform MI300X, analysis mode default, output to
./analysis_output
Adapt the environment fields if TraceLens runs on a remote node or inside a container.
The orchestrator will:
Generate a TraceLens performance report
Prepare per-category data (GEMM, SDPA, norm, etc.)
Launch system-level and compute-kernel sub-agents in parallel
Validate, aggregate, and write
analysis.md
This may take a while. The workflow invokes 13 specialized sub-agents.
Step 4: Review the output#
The only artifact intended for end-user review is analysis.md. Everything
else under the output directory is agent internals.
analysis_output/
├── analysis.md # Stakeholder report (read this)
├── perf_report.xlsx # Internal
├── perf_report_csvs/ # Internal
├── category_data/ # Internal
├── system_findings/ # Internal
├── category_findings/ # Internal
└── metadata/ # Internal
Open analysis.md and check the structure:
Executive Summary: workload characterization, metrics table, improvement chart
Compute Kernel Optimizations: top operations and ranked P-items
Kernel Fusion Opportunities (Experimental)
System-Level Optimizations (Experimental)
Detailed Analysis: per-P-item drill-down with identification, data tables, reasoning, and resolution
Each P-item includes impact_score bounds in HTML comment markers
(<!-- impact-begin ... -->) that downstream optimization tools can parse
programmatically.
Step 5: (Optional) Comparative analysis#
When you have two traces from the same framework (e.g. both vLLM, both SGLang), run a gap analysis. Always pass the baseline trace as trace 1.
Cursor chat:
Follow the analysis orchestrator installed with TraceLens and run the full agentic
analysis workflow on <path_to_baseline_trace.json> and <path_to_primary_trace.json>
Claude Code:
Follow the analysis orchestrator and run the full agentic analysis workflow on
<path_to_baseline_trace.json> and <path_to_primary_trace.json> with platform MI300X,
analysis mode default, output to ./analysis_output_comparative
Comparative output adds per-trace perf reports (perf_report_trace1.xlsx,
perf_report_trace2.xlsx) and gap-based impact estimates in category_data/.
Cross-framework comparisons (for example, vLLM vs. SGLang) might produce misleading gap estimates due to structural differences in operation call stacks.
Step 6: (Optional) Run headless with the Cursor agent CLI#
For batch runs or CI, install the Cursor CLI and pass all parameters inline so no interactive prompts are needed.
curl https://cursor.com/install -fsS | bash
Local, standalone, default mode:
agent --model claude-opus-4-7-high --print --force --trust \
"Follow the analysis orchestrator installed with the TraceLens pip package (look under TraceLens/Agent/Analysis/.cursor/skills/ in the package installation directory) and run the full agentic analysis workflow on <path_to_trace.json> with platform <platform>, analysis mode default, output to <output_dir>"
Cluster + container, inference (vLLM or SGLang eager):
agent --model claude-opus-4-7-high --print --force --trust \
"Follow the analysis orchestrator installed with the TraceLens pip package (look under TraceLens/Agent/Analysis/.cursor/skills/ in the package installation directory) and run the full agentic analysis workflow on <path_to_trace.json> with platform <platform>, analysis mode inference, execution mode eager, node <node>, container <container>, output to <output_dir>"
See
TraceLens Agent Analysis README
for graph replay + capture and additional CLI variants. The eval script
agent_evals/Analysis/eval_scripts/generate_ref.sh in the TraceLens repo is a
reference for batch reference generation.
Step 7: (Optional) Try to get things done without AMD Skills#
Remove the added skill from your agent’s skills directory and rerun the experiment above. Without the orchestrator skill you should see higher variance in execution length, token usage, and report quality. Common issues without the skill include:
Agent improvising ad-hoc trace parsing instead of running TraceLens CLI tools
Missing or inconsistent sub-agent workflow (system-level vs. compute-kernel tiers)
Report that lacks prioritized P-items, impact scores, or the three-tier structure
Skipped validation steps or incomplete
category_data/preparationAgent providing a generic performance essay instead of a structured
analysis.md
Next steps#
Inference profiling: Use the TraceLens profiling skill (
magpie-benchmark-profiling) or try themagpie-kernel-evaluatorwalkthrough to collect vLLM or SGLang traces before analysis.Optimization loop: Parse
analysis.mdP-items and impact markers to drive kernel tuning, fusion, batching, or precision narrowing, then re-profile and re-run analysis to measure improvement.Upstream docs: Full architecture, sub-agent reference, and output layout are documented in the TraceLens Agent Analysis README.