What problem does it solve? Teams routinely ship decisions on analyses that are underpowered, riddled with uncorrected multiple comparisons, or built on unchecked assumptions. This Skill turns each analysis into an explicit verdict — ship, ship with caveats, or do not ship — with the failing check named, so a weak number cannot hide behind a confident headline. ## Core Features & Use Cases - Three mandatory checks: Every review verifies statistical power, multiplicity (all comparisons run, not just reported ones), and method assumptions before any verdict is written. - Verdict with a named reason: Each analysis receives ship, ship with caveats, or do-not-ship, plus the single check that decided it and the most consequential caveat. - Paired-analysis diffing: When two analyses bear on one decision, the review states which one carries the decision and why, or that neither does. - Use Case: Before a rollout go/no-go, hand the Skill two experiment readouts with claims, windows, and sample sizes; it returns a dated markdown file with two verdicts, flagging an underpowered "no effect" as inconclusive rather than a pass. ## Quick Start Ask the assistant to review these two experiment readouts for statistical soundness and give a ship or do-not-ship verdict on each.