What problem does it solve?
Teams running LLM-powered agents lose track of model spend and silently drift into overpriced or outdated models. This Skill performs a recurring, data-grounded audit of model cost and quality so every surface runs a defensible model choice.
Core Features & Use Cases
- Spend and token reporting: Pulls per-surface and per-model cost, token, and request data from the AI Gateway report for two 7-day windows.
- Drift detection: Compares period-over-period cost, model mix, and token shape to flag meaningful changes with reasons.
- Model landscape research: Checks current pricing and quality benchmarks from the AI Gateway catalog, leaderboards, and Artificial Analysis before recommending any swap.
- Use Case: Every Monday, run the watchdog to produce a Linear document with per-surface spend, drift findings, and model recommendations, plus a Linear issue for any decision-worthy change.
Quick Start
Ask the agent to run the weekly cost-watchdog review of model spend and drift for the last full week.