battle

Benchmark AI model configurations on coding tasks to produce a leaderboard.

Updated Dec 17, 2025
One-click install
npx skills add https://github.com/eprouveze/HealthPulse --skill battle
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: battle
Source: https://github.com/eprouveze/HealthPulse/tree/main/.claude/skills/battle
Command: npx skills add https://github.com/eprouveze/HealthPulse --skill battle

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Benchmark AI model configurations on the same coding task to inform model selection.

Core Features & Use Cases

  • Cross-model battles across Claude, Gemini, Codex, Kimi, and DeepSeek on a single coding task to produce a leaderboard.
  • Solo and pair-battle workflows with scoring by tokens, cost, time, and output quality.
  • Repeatable evaluation process to guide model strategy decisions.

Quick Start

Describe the task and start a cross-model battle to generate a leaderboard of results.

Frequently Asked Questions about battle

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark AI models on the same coding task?

Yes, you can compare AI model coding performance by running solo or pair-programming battle workflows. The process applies the same task across Claude, Gemini, Codex, Kimi, and DeepSeek, producing a leaderboard based on cost and quality.

How do I evaluate AI model cost and token usage for coding tasks?

Run a cross-model battle to evaluate AI model cost and token usage. The workflow scores each model on the same coding task, producing a leaderboard that details token consumption, execution time, and output quality.

Does the cross-model AI battle support pair-programming workflows?

The cross-model AI battle supports solo and pair-battle workflows. It evaluates configurations across multiple models on a single coding task, generating a leaderboard based on tokens, cost, time, and output quality.

What is the best way to generate an AI model leaderboard for coding tasks?

Generate an AI model leaderboard by running a cross-model battle on a single coding task. The evaluation process scores configurations across models like Claude, Gemini, Codex, Kimi, and DeepSeek, producing a ranked report based on cost and quality.

Can I benchmark multiple AI model configurations on a single coding task?

Benchmark multiple AI model configurations by running a cross-model battle. The workflow applies the same coding task to each model, producing a final battle report with diff patches and a leaderboard based on token usage and cost.