eval-suite-planner

Generate a structured eval plan from an agent description for MS Learn Stage 1 Define.

123|20|Updated Mar 30, 2026
One-click install
npx skills add https://github.com/microsoft/eval-guide --skill eval-suite-planner
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: eval-suite-planner
Source: https://github.com/microsoft/eval-guide/tree/main/skills/eval-suite-planner
Command: npx skills add https://github.com/microsoft/eval-guide --skill eval-suite-planner

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Produces a complete eval-suite plan from a plain-English agent description, grounding the plan in Microsoft's Eval Scenario Library and MS Learn agent evaluation guidance.

Core Features & Use Cases

  • Defines Step 0 routing to business problem and capability scenario types (Information Retrieval, Knowledge Grounding, etc.) and maps to core plan outputs.
  • Generates a scenario table with core business, capability, edge-case, and variation tests, plus primary and secondary evaluation methods per scenario.
  • Produces accompanying quality signals mapping, pass/fail thresholds, and a rationale that explains how the plan supports Stage 1 Define and enables Stage 2 execution.

Quick Start

Provide an agent description to generate a complete eval-suite plan.

Frequently Asked Questions about eval-suite-planner

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I plan an AI agent evaluation suite from a plain-English description?

Plan an AI agent evaluation suite by providing a plain-English agent description to generate a structured eval plan, mapping to core business, capability, safety, edge-case, and variation scenarios. This grounds the plan in Microsoft's MS Learn agent evaluation guidance.

What is Stage 1 Define planning in the MS Learn four-stage evaluation framework?

Stage 1 Define planning establishes the foundational eval suite structure before execution. It routes business problems to capability scenario types like Information Retrieval and Knowledge Grounding, mapping to core plan outputs including scenario tables and quality signals.

How do I generate test scenarios for Copilot Studio agent evaluation?

Generate test scenarios for Copilot Studio agent evaluation by inputting an agent description. The system produces a scenario table featuring core business, capability, edge-case, and variation tests, alongside primary and secondary evaluation methods per scenario.

Can I use the MS Learn evaluation framework for Information Retrieval and Knowledge Grounding agents?

Yes, the MS Learn evaluation framework supports Information Retrieval and Knowledge Grounding agents by routing Step 0 to specific capability scenario types. It generates quality signals mapping and pass/fail thresholds tailored to these business problems.

What does an eval plan include for AI agent evaluation?

An eval plan for AI agent evaluation includes a detailed scenario table, quality signals mapping, pass/fail thresholds, and planning rationale. It supports optional assets, scripts, and references while satisfying frontmatter constraints.

How do I set pass/fail thresholds for AI evaluation scenarios?

Set pass/fail thresholds for AI evaluation scenarios by generating a structured eval plan from an agent description. The plan automatically produces accompanying quality signals mapping and thresholds that explain how the plan enables Stage 2 execution.