What problem does it solve?
This skill prevents AI agents from producing unreliable research by enforcing hard gates on pre-registration, sample size calculation, and blind independent judging, ensuring conclusions are defensible rather than built on sand.
Core Features & Use Cases
- Pre-registration Lock: Hashes and freezes study designs to prevent HARKing (Hypothesizing After Results are Known).
- Statistical Power Enforcement: Automatically computes required sample sizes to ensure studies are adequately powered.
- Blind Independent Judging: Orchestrates a separate agent to score outputs without bias, blocking conclusions if the judge is unqualified.
- Use Case: Use this when conducting A/B tests, model evaluations, or ablations where the integrity of the conclusion is critical for downstream product decisions.
Quick Start
Use the crucible skill to initialize a new study with the goal of comparing the correctness of two different prompt strategies.