What problem does it solve? Teams often adopt open-weight models based on ideology, leaderboard scores, or vendor-independence vanity, then discover hosting, ops, and review costs erase any savings. This Skill enforces a fail-closed decision process so open-weight or local models are only approved when fully loaded cost per successful task actually beats API pricing under a hard monthly budget cap. ## Core Features & Use Cases - Baseline Measurement: Requires spend, tokens by task, latency, error/retry, and human review data before any model decision. - Fully Loaded TCO: Computes cost per successful task including hosting, ops, and human review, not just token price. - Hybrid Routing Gates: Approves cheap first-pass models only with escalation paths for low-confidence or high-stakes outputs. - Pilot Enforcement: Requires 30-60 day pilots with quality, latency, and savings gates plus an API fallback. - Use Case: Deciding whether to route a repeatable IAP extraction workload through a local open-weight model at $5/month operating cost instead of a paid API, while keeping an API fallback and denying requests to rent GPU fleets under a $20/month cap. ## Quick Start Ask the agent to evaluate whether switching a specific high-volume workload to an open-weight model saves money, providing your monthly API spend, task volumes, and measured quality metrics.