review-llm-annotations-and-improve-prompt

Analyze human versus LLM judge annotation discrepancies to refine evaluation prompts.

2|Updated Feb 17, 2026
One-click install
npx skills add https://github.com/coval-ai/coval-external-skills --skill review-llm-annotations-and-improve-prompt
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: review-llm-annotations-and-improve-prompt
Source: https://github.com/coval-ai/coval-external-skills/tree/main/skills/human-review/review-llm-annotations-and-improve-prompt
Command: npx skills add https://github.com/coval-ai/coval-external-skills --skill review-llm-annotations-and-improve-prompt

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill addresses the gap between machine-generated evaluation scores and human ground truth, helping you identify why an LLM judge is misclassifying data and how to fix its instructions.

Core Features & Use Cases

  • Agreement Analysis: Calculates statistical agreement between human reviewers and LLM judges across binary, categorical, and numerical metrics.
  • Pattern Identification: Analyzes transcripts and reviewer notes to pinpoint systematic biases or edge cases in your evaluation prompts.
  • Prompt Engineering: Provides a structured workflow to draft and apply improved prompts based on real-world disagreement data.

Quick Start

Use the review-llm-annotations-and-improve-prompt skill to analyze project 123 and metric 456 to identify and resolve evaluation discrepancies.

Frequently Asked Questions about review-llm-annotations-and-improve-prompt

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I improve LLM judge prompts when evaluation scores disagree with human annotations?

To improve LLM judge prompts when scores disagree with human annotations, analyze discrepancies between machine-generated scores and ground-truth reviewer feedback to identify systematic biases, then apply structured prompt refinements based on those specific edge cases.

What is the best way to analyze LLM judge agreement across categorical and numerical metrics?

Analyzing LLM judge agreement across categorical and numerical metrics involves calculating statistical alignment between machine-generated scores and human ground truth, aggregating simulation transcripts and reviewer notes to pinpoint systematic misclassifications within evaluation projects.

Why does my LLM judge misclassify data despite detailed evaluation instructions?

An LLM judge misclassifies data despite detailed instructions due to systematic biases or unhandled edge cases. Identifying these patterns requires comparing machine outputs against human ground-truth annotations and refining the evaluation prompt instructions accordingly.

Can I use human reviewer feedback to fix binary LLM evaluation metrics?

Yes, you can use human reviewer feedback to fix binary LLM evaluation metrics by aggregating simulation transcripts and reviewer notes to locate scoring discrepancies, which then generate actionable improvements for the binary evaluation prompt.

Do I need ground-truth annotations to refine LLM evaluation prompts?

Yes, ground-truth annotations are required to refine LLM evaluation prompts. The process systematically compares human reviewer feedback against machine-generated LLM judge scores to identify discrepancies and generate actionable prompt improvements.

What types of LLM judge metrics can be optimized using annotation discrepancies?

Text-based LLM judge metrics including binary, categorical, and numerical types can be optimized using annotation discrepancies by analyzing disagreements between machine scores and human ground truth within Coval evaluation projects.