Pi-Bench
Benchmark for proactive assistants in long-horizon workflows
All Skills in This Repository (29)
Pure Emerald Level Indicatorsmemory
Separate persistent facts from event history and search HISTORY.md with grep.
summarize
Condense long-form text from URLs, local files, and YouTube links into summaries.
clawhub
Search the ClawHub registry by natural-language queries and install skills into the nanobot workspace.
skill-creator
Create and package modular AgentSkills with SKILL.md frontmatter metadata.
github
Retrieve GitHub issues, pull requests, and CI run details via gh CLI.
tmux
Start isolated tmux sessions, send keystrokes, and capture pane output.
weather
Retrieve current weather and forecasts for locations via wttr.in with Open-Meteo JSON fallback.
cron
Schedule reminders and recurring tasks with cron expressions and intervals.
jargon-translator
Translate workplace jargon between plain Chinese and professional phrasing.
law-exam-trainer
Convert law exam videos and documents into structured practice questions.
Data Analysis
Convert raw data into validated insights with uncertainty and caveats.
china-tax-law
Cite PRC tax laws and provisions for compliance and dispute guidance.
Frequently Asked Questions
FAQPage SchemaHow to install Pi-Bench?▼
Run `npx skills add Simplified-Reasoning/Pi-Bench --all -g -y` in your terminal to install all skills in this suite globally.
What does Pi-Bench measure?▼
It measures Proactivity (whether an assistant discovers hidden user intents early) and Completeness (whether final outputs satisfy checklist requirements) across long-horizon, multi-session tasks.
Which models can Pi-Bench evaluate?▼
It ships with ready-made configs for Claude, DeepSeek, Gemini, GLM, MiniMax, and Doubao models, and you can add any custom model via a YAML config file.
How do I run a Pi-Bench evaluation?▼
After setup, run `pibench --model-id deepseek-v3.2 --run 3` from the repository root to execute three leaderboard-style repeats of the benchmark.
What personas and tasks does Pi-Bench include?▼
It includes 100 multi-turn tasks across five personas: researcher, marketer, pharmacist, law trainee, and financier, each with persistent workspaces and hidden intents that emerge over time.
Related Repositories in Data & Analytics
View All in Data & Analytics→PaddleOCR
Extract text, tables, and formulas from PDFs and images
Scrapling
Scrape any website and bypass anti-bot protection with AI
last30days-skill
Research any topic across Reddit, X, YouTube, and the web