ab-test-setup

Design statistically rigorous A/B tests with sample size and duration calculations.

Updated Feb 14, 2026
One-click install
npx skills add https://github.com/lionheartapp/lionheart-ops --skill ab-test-setup-lionheartapp
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ab-test-setup
Source: https://github.com/lionheartapp/lionheart-ops/tree/main/.claude/skills/ab-test-setup
Command: npx skills add https://github.com/lionheartapp/lionheart-ops --skill ab-test-setup-lionheartapp

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill helps teams avoid poorly designed experiments by providing a repeatable framework for forming hypotheses, selecting primary metrics, calculating sample sizes and durations, and running analyses that produce reliable, actionable decisions.

Core Features & Use Cases

  • Hypothesis Framework & Templates: Structured phrasing to turn observations into testable hypotheses and documented test plans.
  • Sample Size & Duration Guidance: Quick reference tables and a duration calculator to estimate required traffic and run time for A/B, A/B/n, and MVT tests.
  • Metric Selection & Guardrails: Guidance on primary, secondary, and guardrail metrics plus analysis checklists and interpretation rules.
  • Use Cases: Optimize homepage CTAs, test pricing page layouts, or iterate on signup flow changes with statistically defensible decisions.

Quick Start

Help me design an A/B test for our pricing page by defining a clear hypothesis, the primary metric, required sample size and duration, traffic allocation, and a rollout checklist.

Frequently Asked Questions about ab-test-setup

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I calculate sample size and duration for an A/B test on a pricing page?

To calculate A/B test sample size and duration, you need baseline conversion rates, minimum detectable effect, traffic allocation, and statistical significance thresholds. This framework provides quick reference tables and calculators to estimate required traffic and run time for reliable product experiments.

What is a guardrail metric and when do I need it for experiment design?

A guardrail metric is a predefined measurement that prevents negative downstream impacts during an experiment. You need guardrail metrics in experiment design to ensure variations like signup flow changes do not degrade critical business metrics while optimizing primary targets.

How do I write a testable hypothesis for a landing page experiment?

Writing a testable hypothesis for a landing page experiment requires structuring observations into clear phrasing that predicts an outcome based on a specific change. This framework provides structured hypothesis templates to turn observations into documented, testable test plans.

Can I use this A/B testing framework for multivariate tests and feature variations?

Yes, you can use this A/B testing framework for multivariate tests and feature variations. It provides sample size and duration calculations specifically supporting A/B, A/B/n, and MVT tests across landing pages, funnels, and various product surfaces.

What statistical significance threshold should I use for product experiments?

Statistical significance thresholds for product experiments depend on your acceptable false positive rate, typically determined by your risk tolerance. This framework provides analysis criteria and interpretation rules to apply consistent significance thresholds across web and product experiments.

Why is my A/B test producing unreliable results despite high traffic?

A/B tests produce unreliable results when sample sizes are too small, durations miss cyclical traffic patterns, or primary metrics lack clear definition. Applying a structured experiment design framework ensures proper metric selection, baseline rates, and statistical rigor for actionable decisions.