exgentic
Standardized evaluation of agents across diverse benchmarks
All Skills in This Repository (2)
Pure Emerald Level IndicatorsFrequently Asked Questions
FAQPage SchemaHow to install Exgentic?โผ
Run `npx skills add Exgentic/exgentic --all -g -y` in your terminal to install all skills in this suite globally.
How to evaluate an agent on a benchmark?โผ
Run `exgentic evaluate --benchmark tau2 --agent tool_calling --subset retail --num-tasks 2 --model gpt-4o` after setting your API key. Benchmarks install automatically on first run.
Which benchmarks and agents does Exgentic support?โผ
It ships with benchmarks like tau2, SWE-bench, AppWorld, BrowseComp, HotpotQA, GSM8K, and BFCL, plus agents including Claude Code, Codex CLI, Gemini CLI, SmolAgents, and LiteLLM tool calling.
Can I add my own agent or benchmark?โผ
Yes. The included add-agent and add-benchmark skills guide you through building thin adapters, registering them, and validating them with smoke tests.
Does Exgentic support isolated evaluation runs?โผ
Yes. Use `--set benchmark.runner=docker` to run evaluations inside Docker containers with only Docker installed locally.
Related Repositories in Data & Analytics
View All in Data & AnalyticsโPaddleOCR
Extract text, tables, and formulas from PDFs and images
Scrapling
Scrape any website and bypass anti-bot protection with AI
last30days-skill
Research any topic across Reddit, X, YouTube, and the web