eval-suite-planning

Plan LLM evaluation test suites from feature specs and architecture guidance.

20|Updated Feb 18, 2026
One-click install
npx skills add https://github.com/Accelerated-Innovation/governed-ai-delivery --skill eval-suite-planning
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: eval-suite-planning
Source: https://github.com/Accelerated-Innovation/governed-ai-delivery/tree/main/agents/codex/skills/backend/eval-suite-planning
Command: npx skills add https://github.com/Accelerated-Innovation/governed-ai-delivery --skill eval-suite-planning

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Plan and orchestrate end-to-end LLM evaluation suites for feature development, coordinating DeepEval, Promptfoo, and RAGAS to ensure governance and quality.

Core Features & Use Cases

  • Determine required evaluation tools from architecture preflight (DeepEval, Promptfoo, RAGAS)
  • Generate eval criteria, dataset references, and test structure for feature evaluation
  • Produce a ready-to-run evaluation plan including artifact layout and thresholds to enforce governance

Quick Start

Provide a complete evaluation plan by reading the feature specs and architecture guidance.

Frequently Asked Questions about eval-suite-planning

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I plan an LLM evaluation test suite for a new feature?

To plan an LLM evaluation test suite, provide feature specs and architecture guidance to determine required evaluation tools, dataset references, test structures, and thresholds for a ready-to-run evaluation plan.

When do I need DeepEval, Promptfoo, or RAGAS for my LLM feature evaluation?

You need DeepEval, Promptfoo, or RAGAS for LLM feature evaluation when your architecture preflight indicates specific testing requirements, which are determined by reading feature specs and architecture guidance.

What's the best way to structure datasets and tests for an LLM evaluation suite?

The best way to structure datasets and tests for an LLM evaluation suite is to generate a recommended tests folder layout that includes dataset references and thresholds based on your feature specs.

Can I generate a ready-to-run evaluation plan that enforces governance and quality thresholds?

Yes, you can generate a ready-to-run evaluation plan that enforces governance and quality thresholds by coordinating DeepEval, Promptfoo, and RAGAS based on architecture preflight and feature specifications.

Does planning an LLM evaluation suite require architecture guidance as input?

Yes, planning an LLM evaluation suite requires architecture guidance as input, along with feature specs, to determine which evaluation tools are required and how to structure datasets and thresholds.