What problem does it solve?
Manually evaluating agent-generated implementation plans for completeness, safety, and alignment with project rules risks accidental workspace changes and inconsistent scoring. This Skill eliminates that risk with a fully non-mutating, evidence-based grading workflow.
Core Features & Use Cases
- Weighted Rubric Scoring: Grades plans on 7 evidence-based dimensions (scope clarity, intent alignment, repo grounding, execution order, harness compatibility, validation, and risk handling) for an objective 0-10 score.
- Multi-Agent Support: Accepts plan files, inline quoted text, or auto-discovers latest plans from Claude Code, Codex, GitHub Copilot, and Gemini CLI.
- Blind Scoring Mode: Strips agent attribution to eliminate bias when comparing plans from different tools.
- Use Case: Engineering leads can quickly validate implementation plans from their team's agentic CLIs to catch missing validation, harness violations, or risky changes before coding begins.
Quick Start
Use the plan-grader skill to score your latest Codex implementation plan and receive a list of blocking gaps to address before running the plan.