What problem does it solve? Verifying a research claim requires testing it against genuinely different methods, datasets, and models, but choosing weak or redundant variants produces verification results that reviewers will not trust. ## Core Features & Use Cases - Candidate Harvesting: Scans existing project research artifacts (idea reports, refinement logs, findings, research wiki) to build a candidate pool before searching externally. - Focused Literature Top-Up: Invokes /research-lit in focused mode only for dimensions with fewer than two credible candidates, avoiding full surveys. - External Reviewer Ranking: Sends the candidate pool to an external LLM reviewer that picks and ranks one swap per dimension by strength of independent test. - Use Case: Given a claim that a probing method detects sentiment in residual streams, the skill selects a different probing submethod, a different-distribution dataset, and a different model family, then writes a structured PLAN.md with expected outcomes and confound controls. ## Quick Start Ask the agent to pick method, dataset, and model swap alternatives for a specific claim ID and statement so the claim can be stress-tested.