calibrate

Calibrate post-launch AI features by documenting error patterns and assessing eval coverage.

16|3|Updated Oct 23, 2025
One-click install
npx skills add https://github.com/breethomas/bette-think --skill calibrate-breethomas
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: calibrate
Source: https://github.com/breethomas/bette-think/tree/main/plugins/bette-think/skills/calibrate
Command: npx skills add https://github.com/breethomas/bette-think --skill calibrate-breethomas

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Provides a repeatable post-launch workflow to document error patterns, assess evaluation effectiveness, and decide whether an AI feature is ready for increased autonomy, reducing regressions and premature promotions.

Core Features & Use Cases

  • Error Pattern Documentation: Guided steps and a template to catalog failures, root causes, fixes, and priorities for triage and tracking.
  • Eval Performance Review: Coverage and gap analysis to ensure evals surface real production issues and evolve with new failure modes.
  • Agency Promotion Decisioning: A checklist-based promotion verdict workflow that evaluates quality, safety, monitoring, and operational readiness.
  • Health Checks & Cadence: Quick checks for weekly health, monthly eval reviews, and quarterly deep calibration cycles for continuous improvement.

Quick Start

Ask calibrate to run a quick health check for feature X and summarize current error patterns, eval gaps, and a promotion recommendation.

Frequently Asked Questions about calibrate

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I monitor AI features post-launch to identify error patterns?

Post-launch monitoring identifies error patterns by cataloging production failures, root causes, and fixes. This Skill provides a structured template to triage and track AI feature errors, ensuring systematic documentation for continuous improvement.

What is the best way to review evaluation effectiveness for production AI features?

Reviewing evaluation effectiveness requires coverage and gap analysis to ensure evals surface real production issues. This Skill assesses whether existing evals evolve with new failure modes and accurately detect post-launch regressions.

How do I decide if an AI feature is ready for agency promotion and increased autonomy?

Deciding on agency promotion requires a checklist-based verdict evaluating quality, safety, monitoring, and operational readiness. This Skill guides promotion decisioning to prevent premature autonomy and reduce regressions in AI-driven product features.

Can I run quick health checks on a specific AI feature after deployment?

You can run quick health checks by requesting a summary of error patterns, eval gaps, and promotion recommendations for a specific feature. This Skill supports weekly health checks, monthly eval reviews, and quarterly deep calibration cycles.

When should I perform an eval gap analysis on my AI product?

Perform eval gap analysis when your AI feature encounters new failure modes or when preparing for agency promotion. This Skill helps determine if your current evals adequately cover production issues and evolve alongside post-launch error patterns.