PawBench
Benchmark models and agent harnesses across 150 tasks
All Skills in This Repository (39)
Pure Emerald Level Indicatorsstate-space-linearization
Calculate Jacobian matrices to linearize nonlinear state-space systems.
mpc-horizon-tuning
Select optimal prediction horizon and cost matrices for MPC tension control.
finite-horizon-lqr
Solve finite-horizon LQR problems for Model Predictive Control using dynamic programming.
integral-action-design
Add integral action to MPC systems for offset-free tension tracking.
manufacturing-failure-reason-codebook-normalization
Normalize engineer-written failure reasons to product codebook standards with confidence scores.
stock-analysis
Analyze historical stock prices and market indicators to identify trends.
file-reading
Read file contents directly for rapid information extraction.
recipe-search
Search recipes by ingredients, cuisine, or dietary restrictions.
music-player
Control music playback via streaming service API integration.
calendar-ops
Automate calendar event management to schedule, update, and remove events.
weather-forecast
Retrieve current and future weather conditions for specified locations via weather APIs.
moltbook-auto-post
Automate content posting and scheduling to the Moltbook platform with rate limiting.
Frequently Asked Questions
FAQPage SchemaHow to install PawBench?โผ
Run `npx skills add agentscope-ai/PawBench --all -g -y` in your terminal to install everything globally. You also need Python 3.11+ and Docker to run evaluations locally.
What does PawBench actually measure?โผ
It measures model and harness performance together across 150 tasks, so you can tell whether a failure comes from the model's reasoning or the agent runtime around it.
How do I compare different agent harnesses?โผ
Fix one model and run the same task set across harnesses like QwenPaw, OpenClaw, and Hermes, then inspect the harness gap and slice diagnostics on the leaderboard.
Which models and harnesses are covered?โผ
Version 1.0 covers 9 models and 3 harnesses (QwenPaw, OpenClaw, Hermes) across 150 tasks tagged by scenario, capability, complexity, modality, and environment.
Can I add my own tasks or harness to PawBench?โผ
Yes. You can contribute new tasks using the five-label taxonomy, add harnesses and graders, and submit run results as JSON files for the public leaderboard.
Related Repositories in Data & Analytics
View All in Data & AnalyticsโPaddleOCR
Extract text, tables, and formulas from PDFs and images
Scrapling
Scrape any website and bypass anti-bot protection with AI
last30days-skill
Research any topic across Reddit, X, YouTube, and the web