What problem does it solve?
Running A/B experiments silently break: sample ratio mismatch biases results, identity fragmentation contaminates variants, exposure stalls mean the experiment measures nothing, and mid-run flag edits rebucket users. This Skill continuously audits PostHog experiments so teams stop shipping decisions on corrupted data.
Core Features & Use Cases
- Validity threat detection: Flags sample ratio mismatch (SRM) via native chi-squared p-values, elevated
$multiple contamination, exposure stalls, mid-run feature flag mutations, and metrics that structurally cannot answer the hypothesis.
- Lifecycle hygiene: Identifies zombie experiments running past their useful life, stopped experiments whose flags still serve multiple variants, and stale drafts, bundled into prioritized recommendations.
- Report authoring with dedupe: Authors or edits inbox reports end-to-end with priorities, suggested reviewers, and chart attachments, using scratchpad memory to avoid duplicate findings across runs.
- Use Case: A team runs a checkout experiment for weeks; the scout detects a 56/44 split on 22k exposures starting at a flag edit, files a P2 report naming the experiment id, flag key, and onset date, and routes it to the experiment owner.
Quick Start
Ask the agent to audit all running PostHog experiments for sample ratio mismatch, contamination, and exposure stalls, and file reports for any confirmed validity threats.