What problem does it solve? Teams routinely call A/B test winners too early, run underpowered tests, or skip hypothesis discipline, producing results that do not replicate. This Skill enforces a predeclared experimentation process so test outcomes are statistically valid and actionable. ## Core Features & Use Cases - Hypothesis and Test Design: Structures hypotheses in a because-we-believe-we-will-know framework, classifies tests as A/B, A/B/n, MVT, or split URL, and defines primary, secondary, and guardrail metrics. - Sample Size and Duration Planning: Provides quick-reference sample size tables by baseline conversion rate and minimum detectable effect, plus duration rules covering day-of-week cycles, B2B business cycles, and sequential testing options. - Experiment Program Management: Supports ICE prioritization, experiment velocity tracking, playbook documentation of winning patterns, and weekly-to-quarterly review cadences. - Use Case: A marketer wants to test a new pricing page headline with 15,000 monthly visitors and a 3.2% signup rate; the Skill calculates the required sample per variant, sets a minimum two-week duration, defines the decision rule, and blocks any early winner call. ## Quick Start Ask the agent to design an A/B test for a specific page change, providing your current conversion rate, monthly traffic, and the smallest improvement worth detecting.