What problem does it solve?
Choosing the right AI model for your gstack skills often relies on guesswork and subjective opinions, leading to wasted time and suboptimal performance. This skill eliminates that guesswork by running the same prompt across multiple AI models and generating hard, comparable data on their real-world performance.
Core Features & Use Cases
- Cross-Model Side-by-Side Testing: Runs identical prompts through Codex, GPT (via Codex CLI), and Gemini to compare their outputs for your specific use case.
- Quantitative Performance Metrics: Measures latency, token usage, cost per run, and optional LLM-judged output quality to give you objective comparison data.
- Use Case: If you are developing a gstack skill for lead generation and want to know whether Codex, GPT, or Gemini delivers the fastest, cheapest, and highest-quality outputs for your skill's prompt, this skill provides the data to make that call.
Quick Start
Invoke the benchmark-models skill to compare Codex, GPT, and Gemini performance on your selected gstack skill prompt.