ab-testing

Design statistically valid A/B tests with sample size and significance calculations.

Updated Mar 17, 2026
One-click install
npx skills add https://github.com/fianchettogianni/my-marketing-skills --skill ab-testing-fianchettogianni
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ab-testing
Source: https://github.com/fianchettogianni/my-marketing-skills/tree/main/skills/ab-testing
Command: npx skills add https://github.com/fianchettogianni/my-marketing-skills --skill ab-testing-fianchettogianni

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill eliminates the risk of running inconclusive, misleading A/B tests that waste marketing and growth resources, by providing proven frameworks for hypothesis writing, sample size calculation, statistical rigor, and systematic experimentation program management.

Core Features & Use Cases

  • Structured Hypothesis Framework: Template to write data-backed, testable hypotheses instead of vague guesses, ensuring every test has a clear expected outcome and success metric.
  • Statistical Rigor Tools: Sample size quick reference tables, significance threshold guidance, and peeking problem warnings to avoid false positives and premature test calls.
  • End-to-End Experimentation Support: Covers individual A/B test design, multivariate test planning, ICE prioritization for experiment backlogs, and playbook documentation to turn test learnings into reusable growth patterns.
  • Use Case Example: A growth marketer testing a homepage headline can use this Skill to calculate required sample size based on their traffic and baseline conversion rate, define primary/secondary/guardrail metrics, and avoid stopping the test early before reaching statistical significance.

Quick Start

Use the ab-testing skill to design a statistically valid A/B test for your pricing page CTA, including a testable hypothesis, required sample size, and clear success criteria.

Frequently Asked Questions about ab-testing

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I calculate the required sample size for an A/B test?

To calculate A/B test sample size, you need your baseline conversion rate and expected effect size. This Skill provides quick reference tables to determine required sample sizes, preventing underpowered test designs and ensuring statistical significance before launching experiments.

How do I write a structured hypothesis for growth experiments?

Writing a structured hypothesis for growth experiments requires a data-backed template rather than vague guesses. This Skill provides a structured hypothesis framework to define clear expected outcomes, testable variables, and primary success metrics for every A/B test.

Why do my A/B tests return inconclusive results and false positives?

A/B tests yield inconclusive results and false positives due to early peeking and underpowered test designs. This Skill mitigates these common pitfalls by enforcing statistical rigor, significance threshold guidance, and warnings against calling tests prematurely.

What is the best way to prioritize an experiment backlog using ICE scoring?

Prioritizing an experiment backlog using ICE scoring involves ranking tests by Impact, Confidence, and Ease. This Skill supports ICE prioritization to systematically rank growth experiments, ensuring your team executes high-value tests before lower-impact ones.

How do I define primary, secondary, and guardrail metrics for split tests?

Defining primary, secondary, and guardrail metrics for split tests ensures you measure intended gains without harming other areas. This Skill supports tiered metric definition to clarify success criteria and monitor unintended consequences during A/B testing.

Can I use this to plan multivariate tests or only individual A/B tests?

You can use this Skill for both individual A/B test planning and multivariate test configuration. It provides end-to-end experimentation support, covering everything from single split tests to scalable experimentation playbook development for growth teams.