What problem does it solve?
It eliminates the friction of planning, executing, and reliably reporting research experiments across compute, sessions, and experiment lifecycle states.
Core Features & Use Cases
- Experiment lifecycle management: draft plans, run self-review, move through pending review/start, and handle revisions.
- Compute and execution orchestration: reserve GPUs, start experiments, run workloads with resilient logging/SSH access, and support parallel sub-agents.
- Progress reporting and durable reporting: heartbeat monitoring for long runs, capture incidents when reusable lessons exist, submit results, and always save a full Markdown experiment report (including uploaded figures via the documents skill).
Use case: You have multiple approved experiment cards and want to start, track progress during long training runs, recover from execution issues, then produce a complete results writeup with charts and artifacts.
Quick Start
Ask your AI agent to run the experiments skill to check your assigned experiments, start the selected pending_start experiment, and then submit results followed by saving the full experiment report.