What problem does it solve?
This Skill eliminates the loss of valuable, reproducible lessons from agent benchmark misses, failed work products, or successful benchmark slices that should be standardized to improve future agentic coding, closed-ended reasoning, and defensive sandbox run performance.
Core Features & Use Cases
- Cross-Benchmark Miss Classification: Categorizes failures from SWE-Bench Pro, HLE-style closed-ended tasks, and defensive ExploitBench runs into actionable, predefined miss classes to pinpoint root causes quickly.
- Narrow Reusable Rule Generation: Converts failure traces and observed success practices into targeted, non-contradictory rules, pruning low-value candidates before promotion to avoid regressions.
- Use Case: After your agent fails a SWE-Bench Pro task due to an implicit contract gap, use this Skill to generate a rule to recover tacit contracts from adjacent code and fixtures before editing, preventing the same failure on similar future tasks.
Quick Start
Use the fairy-tale-benchmark-feedback skill to analyze your recent SWE-Bench Pro miss trace and generate a narrow reusable rule to prevent the same API compatibility break on future similar tasks.