eval-audit

Audit LLM evaluation pipelines for reliability gaps and improvement opportunities.

Updated May 5, 2026
One-click install
npx skills add https://github.com/yanochka11/harness_bro --skill eval-audit-yanochka11
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: eval-audit
Source: https://github.com/yanochka11/harness_bro/tree/main/.claude/skills/curated/evals/eval-audit
Command: npx skills add https://github.com/yanochka11/harness_bro --skill eval-audit-yanochka11

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps teams identify weaknesses in LLM evaluation pipelines, including unreliable metrics, unvalidated judges, missing error analysis, and outdated evaluation practices.

Core Features & Use Cases

  • Eval Pipeline Auditing: Inspect evaluation workflows, artifacts, judges, metrics, and review processes to find high-impact issues.
  • Diagnostic Analysis: Evaluate error analysis, evaluator design, judge validation, human review, labeled data quality, and pipeline hygiene.
  • Use Case: When inheriting an AI application with unclear evaluation reliability, use this Skill to produce a prioritized audit report with concrete improvements.

Quick Start

Use the eval-audit skill to review my LLM evaluation pipeline and identify the most critical reliability problems.

Frequently Asked Questions about eval-audit

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I audit my LLM evaluation pipeline for reliability issues?

To audit an LLM evaluation pipeline, perform diagnostic checks on evaluator quality, judge validation, error analysis, and labeled data quality to identify reliability gaps and produce a prioritized report of concrete improvements.

What is LLM eval judge validation and when do I need it?

LLM eval judge validation is the process of verifying that automated evaluators accurately assess model outputs. You need it when inheriting AI applications with unclear evaluation reliability or deploying new metrics.

How do I find high-impact issues in my AI product evaluation workflow?

Find high-impact issues by inspecting evaluation workflows, artifacts, metrics, and review processes. This pipeline auditing approach highlights missing error analysis and unvalidated judges to target critical weaknesses.

Does this pipeline auditing approach work with existing observability systems and traces?

Yes, pipeline auditing applies directly to AI product evaluation workflows involving traces, judges, metrics, labeled datasets, and observability systems to evaluate infrastructure health and review processes.

What is the best way to improve unvalidated judges in LLM evals?

The best way to improve unvalidated judges is through structured diagnostic analysis that assesses evaluator design and human review processes, yielding a prioritized audit report with concrete improvement opportunities.

What are common limitations when diagnosing LLM evaluation infrastructure health?

Limitations arise when evaluation workflows lack structured artifacts, labeled datasets, or observability systems, making it difficult to perform accurate diagnostic checks on evaluator quality and pipeline hygiene.