What problem does it solve?
This Skill helps you verify whether product-eval’s scoring constants are predicting reality accurately, so you can trust confidence and verdicts instead of guessing when the model feels too strict or too loose.
Core Features & Use Cases
- Outcome-based validation: Replays past decisions against known results and compares predicted verdicts with what actually happened.
- Error diagnosis: Identifies false positives, false negatives, and the dominant calibration bias in the current scoring setup.
- Local tuning recommendations: Suggests one-lever-at-a-time changes to scoring constants and prepares a team-scoped override profile without altering the shipped defaults.
- Use case: A product team with ten or more closed decisions can use this Skill to test whether confidence is overestimating opportunity quality before changing thresholds.
Quick Start
Ask the skill to calibrate your scoring against a set of past decisions with known outcomes and recommend the smallest local constant change that would improve accuracy.