What problem does it solve?
Self-eval provides honest, calibrated work evaluations after completing tasks or reviews by applying a two-axis scoring system and mandatory devil's advocate reasoning, addressing the common bias of inflated self-assessments.
Core Features & Use Cases
- Two-axis scoring: independently rates task ambition and execution quality, then combines via a fixed matrix.
- Mandatory devil's advocate: requires arguing for both higher and lower scores before finalizing.
- Score persistence: appends evaluation results to .self-eval-scores.jsonl, building history across sessions.
- Anti-inflation detection: analyzes past scores to flag clustering and stabilize assessments over time.
- Matrix-based composite scoring: final score derived from a predefined matrix to ensure consistent judgments.
Quick Start
After completing work in a Claude Code session, run /self-eval with context about what was evaluated to generate the assessment.