Darwin-Agent
Official@darwin-agent
Frontier research on Agentic AI: general, evolving, future-facing core technologies for agents.
Agent Skills by Darwin-Agent
Showing 12 vetted skills indexed across 1 GitHub repositories.
gaia-playbook
Diagnose GAIA trajectory gaps and select intervention patterns for score improvements.
tb2-playbook
Define TB2 evolution constraints for sandbox topology, tool constraints, and trajectory signaling.
tau2-playbook
Automate β-bench evaluation setup and domain guidance for HarnessX.
journal
Store cross-round hypotheses and outcomes in YAML frontmatter for meta-agent workflows.
validate
Validate HarnessX artifacts with canonicalization, dry-fire, contract, and synthetic replay checks.
analyze
Analyze task trajectory frontmatter and body excerpts to generate candidate HarnessConfig lever hypotheses.
reference
Document HarnessX component authoring shapes and workflows for tools, processors, and prompts.
xlsx
Process spreadsheet inputs into formatted Excel workbooks with validation.
Extract text and data from PDF documents using Python libraries.
skill-creator
Automate Claude Skill creation, testing, and optimization with eval results.
pptx
Generate, edit, and template PPTX presentations from data sources.
docx
Automate Word document creation, editing, and formatting for .docx outputs.
Frequently Asked Questions About Darwin-Agent
FAQPage SchemaWhat specific tasks are enabled by the Darwin-Agent skill set?âŧ
These skills enable the systematic evaluation of HarnessX benchmarks, the generation of candidate configuration hypotheses, and the processing of unstructured data into structured formats. Users can perform trajectory gap analysis, validate artifact contracts, and generate formatted business documents from raw data inputs.
Which personas benefit most from these evaluation capabilities?âŧ
Research engineers and technical evaluators focused on agentic performance metrics benefit most. These skills are designed for professionals managing complex benchmark topologies, trajectory signaling, and the rigorous validation of synthetic artifacts within experimental research environments.
What are the primary prerequisites for deploying these evaluation skills?âŧ
Deployment requires an existing HarnessX environment and access to the specific playbook modules for GAIA, TB2, or Tau2 benchmarks. Users must ensure their environment supports YAML frontmatter processing for hypothesis tracking and has the necessary dependencies for document manipulation libraries.