model-selection-router

Route AI jobs to cost-appropriate model tiers using classification questions and eval protocols.

Updated Jul 16, 2026
One-click install
npx skills add https://github.com/Cloud-Byte-Consulting/plugins --skill model-selection-router-cloud-byte-consulting
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: model-selection-router
Source: https://github.com/Cloud-Byte-Consulting/plugins/tree/main/ai-operations/skills/model-selection-router
Command: npx skills add https://github.com/Cloud-Byte-Consulting/plugins --skill model-selection-router-cloud-byte-consulting

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve? Teams default to one expensive frontier model for every task, overpaying for routine work while under-specifying genuinely hard jobs. This Skill replaces habit-based routing with a repeatable method: classify the work first, then assign it to a daily-driver, cheap-workhorse, or frontier tier with explicit escalation rules. ## Core Features & Use Cases - Three-tier routing with specialists: Classify jobs using seven questions (ambiguity, definable good, inspectability, sensitivity, modality, action, context location) and assign them to daily driver, cheap workhorse, or frontier tiers, attaching specialist capabilities for vision, live data, or tool use. - Eval protocols and benchmark discipline: Run 30-minute and one-week tests on a personal eval set of 3-5 real tasks, and interpret benchmark launches by per-task results and failure modes rather than leaderboard averages. - Frontier job briefing and launch triage: Spec expensive runs with a nine-field brief (outcome, source pack, tool access, boundaries, review standard, proof trail, human gate) and filter new model launches with a five-question adoption test. - Use Case: Your team pays frontier prices for weekly summary decks. Classify the job, run a one-week challenger test against a cheap model with a four-column log (time, rework, quality, would-send), then set a routing table with escalation triggers and present the evidence to management. ## Quick Start Ask the assistant to classify your recurring AI jobs with the seven questions and build a routing table assigning each job to a daily driver, cheap workhorse, or frontier tier.

Frequently Asked Questions about model-selection-router

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I choose which AI model to use for a task?

Classify the job before naming any model using seven questions covering ambiguity, definable good, inspectability, sensitivity, modality, action, and context location. Familiar tasks with definable good and fast review go to a cheap workhorse; ambiguous or taste-heavy work goes to a daily driver; directionally expensive mistakes justify a frontier model.

How to test whether a cheaper AI model is good enough?

Run a 30-minute test on one recurring artifact through both routes, timing review and labeling outputs usable, repairable, or rejected. Follow with a one-week test of five real artifacts, tracking review minutes, acceptance, and failure mode, and promote the cheap route only where review stays cheap.

Should I switch models when a new one tops the benchmarks?

No, suite leaders can lose specific task types to lower-average rivals, so route by per-task results and failure modes rather than leaderboard rank. Apply a five-question filter covering tool integration, data access, ecosystem, and composability, then run promising candidates against your own eval set.

Does higher reasoning effort always improve model output?

No, maximum reasoning effort has been observed underperforming high effort on long-running work because heavier reasoning burns context and forces compactions. Default to high effort, use extra effort for hard bounded verifiable problems, and test effort settings explicitly as part of model selection.

When should I pay for a frontier model instead of a cheap one?

Use a frontier model only when being directionally wrong is expensive, meaning the shape of the artifact itself is the problem. Brief it with a nine-field spec covering outcome, source pack, tool access, boundaries, review standard, proof trail, and a named human approver.

How do I challenge a bad corporate default AI model?

Log one weekly job through both the default and a challenger for one week using a four-column log of time, rework, quality, and would-you-send. Then ask for one specialist license for one job class, keeping the request smaller than the evidence rather than proposing to replace the default.