What problem does it solve?
Product teams often ship changes without knowing whether they truly moved the metrics that matter, or they misinterpret experiment results due to statistical pitfalls like sample ratio mismatches, false positives, and short-term bias. This Skill provides frameworks from Lenny's Podcast and Newsletter guests (Ronny Kohavi, Archie Abrams, Lauryn Isford, and others) to run rigorous experiments and make sound ship, iterate, or kill decisions.
Core Features & Use Cases
- Hypothesis & Experiment Design: Draft falsifiable hypotheses, choose success and guardrail metrics, and calculate required sample sizes before committing engineering time.
- Statistical Rigor: Apply checks like Sample Ratio Mismatch (SRM), Twyman's Law for suspicious results, CUPED variance reduction, and false positive risk analysis.
- Decision Frameworks: Use long-term holdouts, the Three Reasons to Skip an Experiment framework, and the Kill/Iterate/Ship decision model for failed tests.
- Use Case: A PM wants to test a new onboarding flow. The Skill helps them estimate that detecting a 5% lift on a 10% converting step needs 60,000+ users per variation, suggests lowering the confidence interval to 85%, and defines guardrail metrics to protect long-term retention.
Quick Start
Ask the agent to help design an A/B test for a specific product change, including hypothesis, sample size, guardrail metrics, and validity checks.