What problem does it solve? Running a trustworthy evaluation of the stellar-raven-codemode MCP server requires coordinating routing gates, paid QA batteries, live-data lanes, judge models, budget caps, and failure triage — a process that is easy to get wrong and expensive to redo. This Skill provides a step-by-step runbook that keeps every eval round methodologically sound and cost-controlled. ## Core Features & Use Cases - Instrument selection and gating: Maps each type of change (scoring, catalog, executor, prompt surface) to the right eval lane — routing gate, QA headline sample, agentic lane, plan regrade, or live-data contract — with free preflight checks before any paid spend. - Agent role separation and budget enforcement: Distinguishes the orchestrating agent from spawned answering and judge agents, and enforces fail-closed --max-budget-usd caps, server-revision pins, and remote-identity probes on every paid run. - Verdict review and upstream findings: Requires agentic review of every wrong/partial verdict against live services, root-cause triage, and filing evidence-backed findings in improvements/ — the primary artifact of every round. - Use Case: After changing the search catalog, ask your CLI agent to run evals; it will run the routing gate, launch a budgeted QA sample against a pinned dev server, review judge verdicts, and file upstream gaps. ## Quick Start Ask your agent to run a full eval round on stellar-raven-codemode using the run-evals skill, starting with the free preflight and routing gate.