eval

Define eval criteria and track pass@k metrics for AI development.

24|5|Updated Feb 8, 2026
One-click install
npx skills add https://github.com/Luohaothu/everything-codex --skill eval-luohaothu
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: eval
Source: https://github.com/Luohaothu/everything-codex/tree/main/skills/eval
Command: npx skills add https://github.com/Luohaothu/everything-codex --skill eval-luohaothu

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Evals before coding establish expectations and enable continuous quality checks for AI systems.

Core Features & Use Cases

  • Define capability and regression evals to guide development.
  • Track pass@k metrics and generate reports to monitor regressions.
  • Use in AI product workflows to ship reliable capabilities.

Quick Start

Run an evaluation plan by defining evals and executing checks with the /eval commands.

Frequently Asked Questions about eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define eval criteria before coding to guide AI development?

To define eval criteria before coding, you establish formal evaluation expectations and execute checks using standardized commands. This eval-driven approach enables continuous quality assurance across feature development and model updates.

What is eval-driven development and how does it ensure AI quality?

Eval-driven development is a process where you define capability and regression evaluations before writing code. It ensures AI quality by establishing expectations early and tracking pass@k metrics to monitor for regressions across release cycles.

Does this evaluation framework support regression testing for model updates?

Yes, the evaluation framework supports regression testing for model updates. It tracks pass@k metrics and generates reports to monitor regressions, ensuring reliable capabilities are shipped across release cycles.

Can I use multiple grader types for AI capability evaluations?

Yes, you can use multiple grader types for AI capability evaluations. The framework supports various grader types and standardized workflows to assess and ensure the quality of AI-driven development.

What is the best way to store and manage evaluation data for AI systems?

The best way to store and manage evaluation data is using a formal eval framework with standardized workflows. It provides eval storage and tracks pass@k metrics to maintain continuous quality checks for AI systems.