performance-analysis

Analyze MaxText training runs to identify performance bottlenecks using TSDB, TraceLens, and IRLens.

29|3|Updated Feb 13, 2026
One-click install
npx skills add https://github.com/AMD-AGI/maxtext-slurm --skill performance-analysis-amd-agi
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: performance-analysis
Source: https://github.com/AMD-AGI/maxtext-slurm/tree/main/skills/performance-analysis
Command: npx skills add https://github.com/AMD-AGI/maxtext-slurm --skill performance-analysis-amd-agi

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Post-training performance analysis of MaxText training jobs to identify bottlenecks, efficiency issues, and resource contention across GPU, host, and network layers using tgs_tagger, TraceLens, and IRLens.

Core Features & Use Cases

  • Multi-tool workflow: TSDB-based comparisons, TraceLens performance reports, and IRLens analysis to pinpoint root causes.
  • Actionable results: Generate structured metrics, GPU/utilization breakdowns, and kernel-level insights.
  • Guided steps: Read results, summarize findings, and validate dashboard availability.

Quick Start

Run the analysis workflow on a completed job's artifacts to generate analysis.json and TraceLens reports for review.

Frequently Asked Questions about performance-analysis

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I analyze MaxText training performance to identify GPU bottlenecks?

You can analyze MaxText training performance by running the workflow on completed job artifacts to generate structured metrics, GPU utilization breakdowns, and kernel-level insights using TSDB, TraceLens, and IRLens.

What is the best way to compare performance across multiple MaxText training jobs?

The best way to compare multiple MaxText training jobs is using TSDB-based comparisons within the analysis workflow, which orchestrates single or multi-job diagnostics to pinpoint resource contention and efficiency issues across GPU, host, and network layers.

Can I run performance profiling mid-training, or is it only for post-run diagnostics?

You can run performance profiling mid-training, not just post-run. The workflow applies guided steps to read results, summarize findings, and validate dashboard availability for both mid-training profiling and post-run diagnostics.

Does TraceLens work with IRLens to pinpoint root causes of training efficiency issues?

TraceLens works with IRLens in a multi-tool workflow to pinpoint root causes of training efficiency issues. TraceLens generates performance reports while IRLens provides analysis to detect bottlenecks across GPU, host, and network layers.

Why does my MaxText training run have low GPU utilization?

Low GPU utilization in a MaxText training run can be caused by resource contention across host or network layers. Use the analysis workflow to generate kernel-level insights and actionable metrics that pinpoint the exact bottleneck.