claw-bench

Evaluate AI agents on standardized tasks with automated verification and global leaderboard scoring.

180|21|Updated Mar 14, 2026
One-click install
npx skills add https://github.com/claw-bench/claw-bench --skill claw-bench
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: claw-bench
Source: https://github.com/claw-bench/claw-bench/tree/main/skills
Command: npx skills add https://github.com/claw-bench/claw-bench --skill claw-bench

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires python, pip, git, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill enables you to evaluate your AI agent's performance on the global Claw Bench leaderboard, providing a standardized testbed for real-world task completion.

Core Features & Use Cases

  • Global Benchmarking: Compare your agent's performance against top AI agents worldwide.
  • Task Variety: Assess across 314 tasks in 33 domains and 4 difficulty levels.
  • Automated Verification: Real-time scoring with automated verifiers for fair assessment.
  • Use Case: If you are developing an AI agent for task automation, use this Skill to test its capabilities and see how it ranks against others.

Quick Start

To get started, run the following command to install the Claw Bench task files and verifiers:

pip install git+https://github.com/claw-bench/claw-bench.git

Frequently Asked Questions about claw-bench

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark my AI agent against a global leaderboard?

To benchmark your AI agent, you evaluate its performance on a standardized testbed of real-world tasks, comparing results against a global leaderboard. This Skill installs task files and automated verifiers to score your agent across various domains and difficulty levels.

What is automated verification for AI agent evaluation?

Automated verification for AI agent evaluation is a scoring mechanism that uses automated verifiers to assess task completion in real-time. It ensures fair assessment by standardizing the verification process across 314 tasks in 33 domains and 4 difficulty levels.

How do I install task files and verifiers for agent testing using Python and git?

You install task files and verifiers for agent testing by running a pip install command with the git repository URL. This requires Python, pip, and git to fetch and set up the standardized testbed components for evaluating your AI agent.

Does AI agent benchmarking support different difficulty levels and task domains?

Yes, AI agent benchmarking supports different difficulty levels and task domains. The evaluation testbed includes 314 tasks spanning 33 domains and 4 difficulty levels, allowing you to assess your agent's capabilities across a wide variety of real-world scenarios.

Can I use this benchmarking testbed for an AI agent built for task automation?

Yes, you can use this benchmarking testbed for an AI agent built for task automation. If you are developing an agent for task automation, this Skill tests its capabilities on real-world tasks and shows how it ranks against other agents worldwide.

What are the limitations of using a standardized testbed for evaluating AI agents?

The limitations of using a standardized testbed for evaluating AI agents include being constrained to the 314 tasks and 33 domains provided. While it offers fair automated scoring, performance on this testbed may not fully reflect an agent's capabilities on custom or unstructured real-world tasks outside the specified domains.