eval-agents

Run scenario-based subagent evaluations to identify gaps in AGENTS.md context.

1|2|Updated Jan 16, 2026
One-click install
npx skills add https://github.com/kynetic-ai/kynetic-spec --skill eval-agents
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: eval-agents
Source: https://github.com/kynetic-ai/kynetic-spec/tree/main/.claude/skills/eval-agents
Command: npx skills add https://github.com/kynetic-ai/kynetic-spec --skill eval-agents

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill ensures agents receive complete, safe context by validating AGENTS.md and executing scenario-based subagent evaluations to reveal gaps in guidance and risk of prompt manipulation.

Core Features & Use Cases

  • Spawns structured subagent evaluations to test agents against predefined AGENTS.md scenarios.
  • Grades agent responses to identify gaps in context, safety, and decision logic.
  • Produces actionable recommendations to strengthen documentation, workflows, and agent behavior.

Quick Start

Run the evaluation workflow to validate AGENTS.md against predefined scenarios.

Frequently Asked Questions about eval-agents

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I test if my AGENTS.md provides sufficient context for agents?

Testing AGENTS.md context involves running scenario-based subagent evaluations to grade agent responses against predefined scenarios. This identifies gaps in guidance and outputs a structured report with actionable recommendations.

What is scenario-based agent evaluation and how does it work?

Scenario-based agent evaluation spawns structured subagent tests against predefined AGENTS.md scenarios to grade responses. It reveals gaps in context, safety, and decision logic, producing actionable recommendations to strengthen documentation.

How do I validate agent safety and prevent prompt manipulation in documentation?

Validate agent safety and prevent prompt manipulation by executing scenario-based subagent evaluations against your AGENTS.md. This identifies risk vulnerabilities and generates a structured report with recommended fixes to strengthen safety.

When do I need to run agent evaluations on my workflows?

Run agent evaluations when applying updates to AGENTS.md, skill definitions, or core workflows. This ensures documentation maintains accuracy, completeness, and safety by identifying context gaps before deployment.

Does agent evaluation work without any external dependencies?

Agent evaluation works without external dependencies, relying on spawning internal subagents to test scenarios. It validates AGENTS.md and skill definitions natively without requiring additional setup.