calibrate

Validate product-eval scoring constants against known decision outcomes.

14|Updated Jun 29, 2026
One-click install
npx skills add https://github.com/sparkline-ventures/product-eval --skill calibrate-sparkline-ventures
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: calibrate
Source: https://github.com/sparkline-ventures/product-eval/tree/main/skills/calibrate
Command: npx skills add https://github.com/sparkline-ventures/product-eval --skill calibrate-sparkline-ventures

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill helps you verify whether product-eval’s scoring constants are predicting reality accurately, so you can trust confidence and verdicts instead of guessing when the model feels too strict or too loose.

Core Features & Use Cases

  • Outcome-based validation: Replays past decisions against known results and compares predicted verdicts with what actually happened.
  • Error diagnosis: Identifies false positives, false negatives, and the dominant calibration bias in the current scoring setup.
  • Local tuning recommendations: Suggests one-lever-at-a-time changes to scoring constants and prepares a team-scoped override profile without altering the shipped defaults.
  • Use case: A product team with ten or more closed decisions can use this Skill to test whether confidence is overestimating opportunity quality before changing thresholds.

Quick Start

Ask the skill to calibrate your scoring against a set of past decisions with known outcomes and recommend the smallest local constant change that would improve accuracy.

Frequently Asked Questions about calibrate

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I validate scoring constants against known decision outcomes?

Validating scoring constants against known decision outcomes requires replaying historical cases to recompute evidence weight, confidence, and verdicts, then comparing predicted outcomes with actual results to identify calibration bias and false positives.

Why does my product evaluation confidence feel too hot or too cold?

Product evaluation confidence feels too hot or too cold when scoring constants are misaligned with reality, causing false positives or false negatives that require threshold validation and confidence tuning against historical decision data.

How do I tune confidence thresholds using historical product decisions?

Tuning confidence thresholds using historical product decisions involves applying one-lever-at-a-time changes to scoring constants, generating local tuning recommendations that prepare a team-scoped override profile without altering shipped default scoring rules.

Do I need a minimum number of closed decisions to run decision analysis calibration?

Decision analysis calibration requires a minimum of ten closed product decisions with known outcomes to effectively test whether confidence is overestimating opportunity quality before applying changes to scoring thresholds.

Can I adjust product evaluation scoring without altering the default team profile?

Adjusting product evaluation scoring without altering the default profile is possible by preparing a team-scoped override profile that keeps recommended calibration changes local, ensuring shipped default scoring constants remain unchanged.

What is the best way to diagnose false positives in product evaluation scoring?

Diagnosing false positives in product evaluation scoring is best achieved by running outcome-based validation that replays past decisions against known results, identifying the dominant calibration bias in the current scoring setup.