What problem does it solve?
This Skill streamlines the complex, multi-step process of running SWE-bench evaluations, managing containerized environments, and analyzing agent-generated patches.
Core Features & Use Cases
- Automated Pipeline: Orchestrates the full lifecycle from container preparation and agent execution to patch extraction and test evaluation.
- Artifact Management: Provides structured logging and audit trails for every test run, including stdout, context files, and patch diffs.
- Use Case: Use this to run a batch of GitHub issues through an AI agent, automatically verify the generated fixes against the project test suite, and generate a summary report of pass/fail results.
Quick Start
Use the swe_bench skill to start the evaluation pipeline for the next five test cases using the cuda llm group.