What problem does it solve?
This Skill allows users to measure the impact of Claude Code rules, skills, and workflows by running quantitative before/after comparisons and evaluating the outputs against assertions.
Core Features & Use Cases
- Benchmark and Evaluate: Measure the impact of rules, skills, and workflows through quantitative comparisons.
- Assertion-Based Evaluation: Grade outputs against falsifiable, discriminative assertions.
- Isolation and Robustness: Ensures filesystem sandboxing and global rule auto-hiding for accurate evaluation.
- Customizable Workflows: Tailor benchmarking procedures to different target types and scenarios.
- Data Analysis and Reporting: Generate detailed reports with pass rates, timing, and evidence of evaluation.
Quick Start
Run a benchmark to compare the performance of a rule pack with 'claude benchmark <rule-pack-name> --with rules <rule-pack-path> --without rules ~/.claude/rules/'.