experiment-metrics

Score candidate experiment metrics using the STEDII framework.

20|4|Updated Oct 4, 2025
One-click install
npx skills add https://github.com/coalesce-labs/catalyst --skill experiment-metrics
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: experiment-metrics
Source: https://github.com/coalesce-labs/catalyst/tree/main/plugins/pm/skills/experiment-metrics
Command: npx skills add https://github.com/coalesce-labs/catalyst --skill experiment-metrics

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

STEDII framework helps teams select trustworthy experiment metrics, ensuring metric validity and reliability to guide data-driven decisions.

Core Features & Use Cases

  • Defines a primary metric and 3-5 guardrail metrics for experiments.
  • Provides a six-dimension STEDII scoring rubric (Sensitive, Timely, Efficient, Debuggable, Interpretable, Isolated) to evaluate candidate metrics.
  • Includes pre-experiment checks (A/A sanity checks, variance assessment, sample size planning) and guidance for segmentation planning.
  • Offers a structured decision framework and practical examples to apply metrics decisions in real projects.

Quick Start

Define a primary metric and 3-5 guardrail metrics for your upcoming experiment using the STEDII framework and validate readiness with a pre-experiment checklist.

Frequently Asked Questions about experiment-metrics

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I select trustworthy metrics for A/B testing?

Select trustworthy A/B testing metrics by evaluating candidates against the STEDII framework, which scores sensitivity, timeliness, efficiency, debuggability, interpretability, and isolation to ensure metric validity and guide data-driven decisions.

What are guardrail metrics and how many should I define for an experiment?

Guardrail metrics protect against unintended business harm during experiments. You should define one primary metric and 3-5 guardrail metrics, governed by a structured decision framework to validate readiness before running product experiments.

How do I estimate sample size and run pre-experiment checks for AB testing?

Estimate sample size and run pre-experiment checks by conducting A/A sanity tests, assessing metric variance, and planning segmentation. These statistical power checks validate metric reliability before launching your experiment.

What is the STEDII framework for experiment metrics evaluation?

The STEDII framework is a six-dimension rubric—Sensitive, Timely, Efficient, Debuggable, Interpretable, and Isolated—used to score candidate metrics during experiment planning, ensuring selected metrics are valid, reliable, and measurable.

When do I need a metrics framework for product experiments?

You need a metrics framework when planning product experiments to score candidate metrics, define primary and guardrail metrics, run pre-experiment variance assessments, and establish governance to ensure trustworthy data-driven decisions.