What problem does it solve?
Choosing the right AI model for your work is usually based on hype or anecdotes instead of real performance data. This Skill removes the guesswork by running the same prompt across multiple LLM providers and measuring objective, comparable metrics.
Core Features & Use Cases
- Cross-model side-by-side comparison: Run identical prompts through Claude, GPT, and Gemini to compare latency, token usage, and cost in a single table.
- Optional quality scoring: Use an LLM judge to rate each model's output on a 0-10 scale to measure output quality, not just speed and cost.
- Baseline tracking: Save benchmark results as JSON to compare future runs and catch model performance regressions over time.
- Use Case: If you are deciding which model to use for a new coding skill, run this benchmark to see which provider delivers the best balance of speed, cost, and output quality for your specific prompt.
Quick Start
Use the benchmark-models skill to compare the performance of Claude, GPT, and Gemini on your specific task prompt to identify the best model for your needs.