What problem does it solve?
This Skill eliminates the risk of running flawed A/B tests that produce false positives, p-hacked results, and incorrect product decisions, ensuring your experiments yield statistically valid, actionable insights.
Core Features & Use Cases
- End-to-end experiment design: Guides hypothesis formation, pre-registration, metric selection (overall evaluation criterion, guardrail metrics), and sample size/power calculation to avoid underpowered or biased tests.
- Robust statistical analysis: Supports both frequentist and Bayesian methods, includes multiple comparison correction, confidence interval reporting, and variance reduction techniques like CUPED.
- Pitfall mitigation: Detects and avoids common failures including peeking, survivorship bias, Simpson's paradox, Sample Ratio Mismatch, and interference effects.
- Use case: A product team testing a new checkout flow can use this Skill to design a valid experiment, calculate the required sample size, analyze results correctly, and avoid shipping a change that harms conversion.
Quick Start
Use the ab-testing skill to design a valid A/B test for the new checkout flow, calculate the required sample size, and analyze the results for statistical and practical significance.