What problem does it solve?
This Skill helps product teams make reliable A/B test decisions instead of guessing from raw conversion numbers. It validates whether an experiment was run correctly, checks if the result is statistically and practically meaningful, and prevents bad launches when SRM, weak sample size, or guardrail regressions make the evidence unreliable.
Core Features & Use Cases
- Experiment validity checks: Verifies sample size, runtime coverage, randomization quality, SRM, and novelty effects before trusting any result.
- Decision-ready statistical analysis: Calculates conversion rates, lift, p-value, and 95% confidence intervals, then compares the outcome against PMContext success thresholds.
- Guardrail-aware launch recommendations: Reviews downside metrics such as revenue, retention, or latency and downgrades decisions when the primary metric improves but important safeguards worsen.
- Use cases: Analyze onboarding experiments, pricing tests, feature rollout experiments, activation funnel changes, or any split test where product managers need a clear ship, extend, or stop recommendation tied back to documented product metrics.
Quick Start
Ask the pm-abtest skill to analyze your A/B test by providing control and variant sample sizes, conversions, runtime, and the relevant PMContext success threshold and guardrail metrics.