eval-audit

Audit LLM evaluation pipelines and prioritize systemic issues in error analysis and judge design.

Updated May 5, 2026
One-click install
npx skills add https://github.com/iani-kuli/harness_bro --skill eval-audit-iani-kuli
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: eval-audit
Source: https://github.com/iani-kuli/harness_bro/tree/main/.claude/skills/curated/evals/eval-audit
Command: npx skills add https://github.com/iani-kuli/harness_bro --skill eval-audit-iani-kuli

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill identifies gaps and weaknesses in LLM evaluation pipelines, such as unvalidated judges, vanity metrics, or missing error analysis, ensuring your evaluation infrastructure is actually trustworthy.

Core Features & Use Cases

  • Diagnostic Auditing: Systematically inspects traces, judge prompts, and metrics across six critical areas of evaluation hygiene.
  • Prioritized Findings: Generates a report ordered by impact, linking each identified problem to a concrete, actionable fix.
  • Use Case: Use this when inheriting an existing evaluation system or when you suspect your current metrics are not accurately reflecting real-world model performance.

Quick Start

Run the eval-audit skill to inspect the current observability data and provide a prioritized list of improvements for our evaluation pipeline.

Frequently Asked Questions about eval-audit

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I audit LLM evaluation pipelines for reliability issues?

To audit LLM evaluation pipelines, systematically inspect observability traces, judge prompts, and metrics to identify systemic issues in error analysis, judge design, and validation processes. This diagnostic check prioritizes findings by impact to ensure production deployment readiness.

Why does my LLM evaluation metrics report not match real-world model performance?

Your LLM evaluation metrics may not match real-world performance due to systemic pipeline issues like unvalidated judges, vanity metrics, or missing error analysis. Auditing observability traces and evaluator configurations helps identify and prioritize these specific reliability gaps.

What do I need to validate LLM judges and fix vanity metrics?

To validate LLM judges and fix vanity metrics, you need access to observability traces, evaluator configurations, and human-labeled datasets. These inputs enable a diagnostic audit across six critical areas of evaluation hygiene to generate a prioritized list of actionable fixes.

Can I use an audit to fix an inherited LLM evaluation system?

Yes, you can use an audit to fix an inherited LLM evaluation system by systematically inspecting existing traces, judge prompts, and metrics. It generates a prioritized report ordered by impact, linking each identified problem to a concrete, actionable fix.

What is the best way to diagnose missing error analysis in AI product development?

The best way to diagnose missing error analysis in AI product development is to perform a systematic audit of LLM evaluation pipelines. This process inspects observability traces and evaluator configurations to identify, prioritize, and provide actionable fixes for evaluation hygiene gaps.