SKYLENAGE-AISKYLENAGE-AIOfficialยท7 Agent Skills Included

QwenClawBench

Realistic benchmark for evaluating OpenClaw agent performance

Evaluates OpenClaw agents against 100 realistic tasks across 8 domains, from workflow orchestration to finance and security. Runs each task in an isolated Docker container with parallel execution, anomaly detection, and resumable runs. Combines automated checks with LLM judge scoring to deliver trustworthy, reproducible agent performance results.
npx skills add SKYLENAGE-AI/QwenClawBench --all -g -y

All Skills in This Repository (7)

Pure Emerald Level Indicators

Frequently Asked Questions

FAQPage Schema
How to install QwenClawBench?โ–ผ

Run `npx skills add SKYLENAGE-AI/QwenClawBench --all -g -y` in your terminal to install everything globally.

What does QwenClawBench measure?โ–ผ

It measures how well OpenClaw agents complete 100 realistic tasks across 8 domains, including workflow orchestration, system administration, finance, data analysis, and security.

How does QwenClawBench score agent performance?โ–ผ

It combines deterministic automated checks with an LLM judge in a hybrid mode, and zeroes out judge scores when basic deliverable checks fail to prevent inflated results.

Can I resume an interrupted benchmark run?โ–ผ

Yes. Re-run the same command with the same output directory to skip completed tasks, or add --rerun-anomalous to retry only failed tasks.

What do I need to run QwenClawBench?โ–ผ

You need Python 3.10 or later, Docker, the OpenClaw Docker image, and API credentials configured in the openclaw_config folder.

Related Repositories in Software Engineering

View All in Software Engineeringโ†’