benchmark-models

Compare Claude, GPT, and Gemini model latency, tokens, and quality.

Updated Jun 9, 2026
One-click install
npx skills add https://github.com/ericdahl-dev/coauthor-cleaner --skill benchmark-models-ericdahl-dev
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: benchmark-models
Source: https://github.com/ericdahl-dev/coauthor-cleaner/tree/main/.claude/skills/gstack/benchmark-models
Command: npx skills add https://github.com/ericdahl-dev/coauthor-cleaner --skill benchmark-models-ericdahl-dev

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires Bash, Read, AskUserQuestion, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This skill allows you to benchmark and compare the performance of Claude, GPT, and Gemini AI models for a specific prompt or gstack skill.

Core Features & Use Cases

  • Cross-Model Benchmarking: Simultaneously runs the same prompt through Claude, GPT, and Gemini and compares the results in terms of latency, tokens, and optionally quality.
  • Quality Analysis: Utilizes the LLM judge tool for optional quality evaluation, providing insights beyond raw latency and token output.
  • Model Comparison: Assesses which model performs best for the given prompt, based on the comparison of performance metrics.

Quick Start

To initiate the benchmark, provide a prompt or choose a gstack skill for comparison using the appropriate command.

Frequently Asked Questions about benchmark-models

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I compare AI model performance across Claude, GPT, and Gemini?

You can compare AI model performance by running the same prompt through Claude, GPT, and Gemini simultaneously. The benchmark measures latency, token usage, and optionally quality to determine which model performs best for your specific prompt.

What is LLM evaluation and how does an LLM judge work for quality analysis?

LLM evaluation assesses model output quality beyond raw metrics like latency and token count. The LLM judge tool analyzes generated responses to provide detailed insights into comparative quality across Claude, GPT, and Gemini results.

Do I need API authorization for model providers to run cross-model benchmarks?

Yes, running cross-model benchmarks requires configuration and authorization for model providers. You must set up access credentials for Claude, GPT, and Gemini APIs before initiating the performance comparison.

Can I use a gstack skill for AI benchmarking instead of a custom prompt?

Yes, you can initiate the benchmark by choosing a gstack skill for comparison instead of providing a custom prompt. The tool evaluates the skill across Claude, GPT, and Gemini to assess performance metrics.

What metrics are measured during an AI benchmark of Claude, GPT, and Gemini?

The benchmark measures latency, token usage, and optionally output quality. It provides detailed insights into these performance metrics to help assess which model performs best for your given prompt.