eval-audit

Audit LLM evaluation pipelines and deliver prioritized remediation findings.

5|Updated Oct 22, 2025
One-click install
npx skills add https://github.com/marchatton/agent-skills --skill eval-audit
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: eval-audit
Source: https://github.com/marchatton/agent-skills/tree/main/.agents/skills/08-evals/eval-audit
Command: npx skills add https://github.com/marchatton/agent-skills --skill eval-audit

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill identifies critical flaws in your LLM evaluation pipelines, ensuring your metrics are trustworthy and your AI product is genuinely improving.

Core Features & Use Cases

  • Diagnostic Checks: Systematically inspects six key areas of your eval pipeline, including error analysis, evaluator design, judge validation, human review, labeled data, and pipeline hygiene.
  • Prioritized Findings: Delivers a report of identified problems, ordered by their impact on your product's success.
  • Concrete Next Steps: Provides actionable recommendations, often referencing other skills, to fix identified issues.
  • Use Case: You've inherited an LLM evaluation system and are unsure if the reported metrics accurately reflect performance. This Skill will audit the existing setup and provide a clear roadmap for improving its reliability and trustworthiness.

Quick Start

Audit the current LLM evaluation pipeline for potential issues and provide a prioritized list of findings.

Frequently Asked Questions about eval-audit

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I audit LLM evaluation pipelines for missing error analysis?

Auditing LLM evaluation pipelines systematically inspects six key areas including error analysis to identify critical problems. It checks eval artifacts like traces, evaluator configs, and metrics dashboards to deliver a prioritized report of findings with concrete next steps for remediation.

What are common problems with LLM evals that use unvalidated judges?

Common problems with unvalidated judges include vanity metrics and inadequate human review processes that fail to reflect genuine AI performance. An LLM eval audit inspects judge validation and evaluator design to identify these flaws and provide actionable recommendations for improving reliability.

Can I audit my LLM evaluation setup if I don't have existing metrics dashboards?

Yes, you can audit your LLM evaluation setup without existing metrics dashboards. The audit can guide you through setting up a new eval infrastructure from scratch, systematically inspecting pipeline hygiene, labeled data, and human review processes to ensure reliable product improvement tracking.

How do I check if my LLM evaluation metrics are trustworthy?

Checking LLM evaluation metric trustworthiness requires systematically inspecting evaluator design, judge validation, and pipeline hygiene. An audit reviews eval artifacts like traces and labeled data to identify vanity metrics and delivers a prioritized list of findings for remediation.

What's the best way to find flaws in an inherited LLM evaluation system?

The best way to find flaws in an inherited LLM evaluation system is running a diagnostic audit across six key areas including error analysis and human review. This inspects existing eval artifacts to produce a prioritized findings report with concrete next steps for improving reliability.