What problem does it solve? Planning a usability test involves defensible decisions about sample size, study mode, recruiting, and piloting that most teams get wrong — reporting fake percentages from 5 users, running one big study instead of iterating, or skipping the pilot. This Skill turns a research question into a complete, defensible test plan grounded in NN/g methodology. ## Core Features & Use Cases - Qual/quant fork and sample sizing: Decides between formative (5 users, 3–4 per divergent segment) and summative studies (~40 for task metrics, ~39 for eyetracking, ~15 for card sorts), with the discovery-curve math to defend small N. - Mode selection matrix: Chooses among moderated/unmoderated × remote/in-person with explicit validity rules, such as requiring moderation for fragile prototypes or "why" questions. - Mandatory pilot with fix-then-freeze: Enforces a dress-rehearsal pilot at least 2 days before the first session and locks the protocol afterward so sessions stay comparable. - Use Case: A designer needs to validate a beta booking flow before release. The Skill produces an 8-section plan: research questions, qualitative study type with an explicit "no metrics claimed" line, 5 participants plus one floater, moderated remote mode with rejected alternatives, task count, logistics, a scheduled pilot date, and stated limitations. ## Quick Start Plan a usability test for our checkout flow prototype and tell me how many participants I need and whether to run it moderated or unmoderated.