ab-test-setup

Plan statistically valid A/B tests with sample size and duration guidance.

2|Updated Mar 20, 2026
One-click install
npx skills add https://github.com/multiplex-ai/muggle-ai-teams --skill ab-test-setup-multiplex-ai
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ab-test-setup
Source: https://github.com/multiplex-ai/muggle-ai-teams/tree/main/skills/ab-test-setup
Command: npx skills add https://github.com/multiplex-ai/muggle-ai-teams --skill ab-test-setup-multiplex-ai

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

A/B testing design and execution framework that helps teams plan experiments, frame hypotheses, and determine reliable results with clear decision criteria.

Core Features & Use Cases

  • Hypothesis-driven test design and framing
  • Supports A/B, A/B/n, and multivariate (MVT) testing with guidance on sample size and duration
  • Primary, secondary, and guardrail metrics to ensure meaningful and safe decisions
  • Use cases across landing pages, pricing, signup flows, and feature changes

Quick Start

Start by drafting a test hypothesis and sample size plan using the provided templates.

Frequently Asked Questions about ab-test-setup

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I calculate sample size and duration for an A/B test?

A/B test sample size and duration are determined by classifying primary, secondary, and guardrail metrics, then applying sample size calculation and sequential testing guidance to achieve statistically significant lift estimates.

What is the best way to frame a hypothesis for landing page experimentation?

Framing a hypothesis for landing page experimentation requires defining the expected change, classifying primary and guardrail metrics, and establishing clear decision criteria to ensure reliable lift estimates for variant comparison.

Can I run multivariate testing and A/B/n experiments for pricing page changes?

Multivariate testing and A/B/n experiments for pricing page changes are supported through guidance on sample size calculation, duration, and sequential testing options to compare multiple variants and obtain reliable lift estimates.

What are guardrail metrics and why do I need them for feature change experiments?

Guardrail metrics protect feature change experiments by monitoring safety indicators alongside primary and secondary metrics, ensuring meaningful and safe decisions without unintended negative impacts on user experience.

When should I use sequential testing options instead of standard A/B testing?

Use sequential testing options instead of standard A/B testing when continuous evaluation and early stopping are needed, applying statistical significance checks to maintain reliable lift estimates without inflating false positive rates.