training-run-triage

Analyze machine learning training runs and generate triage reports with recommendations.

1|Updated Apr 12, 2026
One-click install
npx skills add https://github.com/KirillKlem/codex-skills --skill training-run-triage
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: training-run-triage
Source: https://github.com/KirillKlem/codex-skills/tree/main/skills/training-run-triage
Command: npx skills add https://github.com/KirillKlem/codex-skills --skill training-run-triage

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides a comprehensive analysis of machine learning training runs, identifying issues such as instability, inefficiency, and degradation, and recommending next steps for improvement.

Core Features & Use Cases

  • Training Run Analysis: Analyze logs, metrics, artifacts, config, and system telemetry to diagnose issues in training runs.
  • Triage Report: Generate a detailed report with findings, recommendations, and evidence-based insights.
  • Use Case: When a user encounters a problematic training run, this Skill can be used to understand the root cause and suggest solutions.

Quick Start

Analyze the training run logs for the experiment with ID '12345'.

Frequently Asked Questions about training-run-triage

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I diagnose machine learning training run instability and performance degradation?

To diagnose machine learning training run instability and performance degradation, you need to analyze training logs, metrics, artifacts, and system telemetry to identify inefficiencies and generate an evidence-based triage report with recommendations.

What is the best way to triage a problematic ML training run?

The best way to triage a problematic ML training run is to analyze experiment logs, configuration files, and system telemetry to identify the root cause of quality issues and suggest concrete next steps for improvement.

Can I use training metrics and artifacts to find the root cause of model training failure?

Yes, you can use training metrics and artifacts to find the root cause of model training failure by analyzing them alongside system telemetry to detect instability, inefficiency, and degradation patterns.

Do I need system telemetry to analyze training run quality issues?

Yes, you need system telemetry to analyze training run quality issues effectively, as it provides the necessary hardware and system-level context to identify inefficiencies and degradation in your machine learning experiments.

What does a training run triage report include?

A training run triage report includes detailed findings on instability, inefficiency, and degradation, along with evidence-based insights and recommendations for improving your machine learning training runs.

Why does my machine learning training run suffer from performance inefficiency?

Your machine learning training run may suffer from performance inefficiency due to configuration issues, system-level bottlenecks, or unstable metrics, which can be identified by analyzing telemetry and artifacts to produce an evidence-based diagnosis.