eval-audit

Audit LLM eval pipelines across six diagnostic areas and output a prioritized remediation report.

Updated Mar 17, 2026
One-click install
npx skills add https://github.com/Avi977/ace-claude-toolkit --skill eval-audit-avi977
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: eval-audit
Source: https://github.com/Avi977/ace-claude-toolkit/tree/main/skills/eval-audit
Command: npx skills add https://github.com/Avi977/ace-claude-toolkit --skill eval-audit-avi977

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Audit and identify gaps in LLM eval pipelines, including missing error analysis, unvalidated judges, vanity metrics, and weak governance.

Core Features & Use Cases

  • Diagnostic checks across six areas: error analysis, evaluator design, judge validation, data labeling, pipeline hygiene, and human review. It prioritizes findings by impact and outputs a structured remediation plan. It also includes a scenario for when there is no eval infrastructure and guidance on how to proceed.

Quick Start

Run an end-to-end audit on your eval pipeline by loading traces, evaluator configs, judge prompts, and labeled data to produce a prioritized findings report.

Frequently Asked Questions about eval-audit

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I audit an LLM eval pipeline for missing error analysis and weak judge validation?

To audit an LLM eval pipeline, run end-to-end diagnostic checks across error analysis, evaluator design, judge validation, data labeling, pipeline hygiene, and human review. This produces a prioritized findings report with structured remediation steps.

What is the best way to identify vanity metrics and unvalidated judges in my LLM evaluation system?

The best way to identify vanity metrics and unvalidated judges is applying structured diagnostic checks to your evaluator configs and judge prompts. This process prioritizes discovered gaps by impact and outputs a remediation-oriented governance report.

How do I validate existing evaluators and improve data hygiene in my AI experiments?

You validate existing evaluators and improve data hygiene by loading your traces, evaluator configs, and labeled data into an audit process. It enforces structured diagnostic checks and outputs a prioritized remediation plan for your AI experiments.

Can I audit my data quality and eval pipeline if I have no existing eval infrastructure?

Yes, you can audit data quality without existing infrastructure. The audit includes a specific scenario for missing eval infrastructure, providing structured guidance on how to proceed and build trustworthy evaluation systems from scratch.

What does an LLM eval audit cover when checking evaluator design and human review processes?

An LLM eval audit covers six core areas: error analysis, evaluator design, judge validation, data labeling, pipeline hygiene, and human review. It enforces diagnostic checks across these areas and prioritizes findings by impact for remediation.