What problem does it solve?
This Skill helps you watch active RL training runs for early warning signs so you can catch instability, resource issues, and metric regressions before they waste time or corrupt a study.
Core Features & Use Cases
- Run Health Monitoring: Checks process state, GPU assignment, queue status, and logging behavior during live training.
- Failure Detection: Flags NaN/Inf events, simulator crashes, episode collapse, reward domination, saturation, and missing metrics.
- Operational Guidance: Produces a clear continue, wait, or stop-investigate recommendation without exceeding contract authority.
- Use Case: Use it when a training job is running but the logs look suspicious, GPU utilization drops unexpectedly, or a guardrail appears to be regressing.
Quick Start
Ask the assistant to monitor the active RL run using the run ID, logs, GPU status, metric stream, and guardrails, then return a concise health summary and recommendation.