hamel-husain

Build production-grade eval sets with human-validated LLM-as-judge workflows.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/voidborne-d/master-skill --skill hamel-husain
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: hamel-husain
Source: https://github.com/voidborne-d/master-skill/tree/main/prototypes/monetize-agents-master/output/sub-skills/hamel-husain
Command: npx skills add https://github.com/voidborne-d/master-skill --skill hamel-husain

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Identify and close AI agent reliability gaps by building production-grade eval sets before making prompts or model changes.

Core Features & Use Cases

  • Evals-first workflow: manual trace review, human-validated rubrics, and LLM-as-judge alignment.
  • Role-play and identity guidance: define role-specific agent behavior and governance.
  • Production tracing: instrumentation and measurement to support continuous improvement.

Quick Start

Identify your current agent's evals gap and align on building an evals-first engagement plan.

Frequently Asked Questions about hamel-husain

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I build production-grade LLM eval sets before changing prompts or models?

Production-grade LLM eval sets require a domain expert to manually annotate 50 traces, implement an LLM-as-judge with human validation, and establish instrumentation for continuous trace collection and evaluation.

What is an evals-first workflow for AI agents?

An evals-first workflow prioritizes manual trace review, human-validated rubrics, and LLM-as-judge alignment to identify and close AI agent reliability gaps before making any prompt or model modifications.

How do I validate an LLM-as-judge against human annotations?

Validating an LLM-as-judge involves having a domain expert manually annotate traces to create a baseline, then aligning the automated judge scoring against these human-validated rubrics to ensure measurement accuracy.

How many manually annotated traces do I need for reliable LLM evaluation?

Reliable LLM evaluation requires a domain expert to manually annotate 50 production traces to build a human-validated baseline for establishing accurate LLM-as-judge rubrics.

Does setting up AI agent tracing require production instrumentation?

Yes, production tracing requires instrumentation to measure agent behavior, collect traces continuously, and support ongoing evaluation to guide system design and model improvements across different domains.

When should I not use an evals-first approach for AI projects?

An evals-first approach is not suitable for AI projects lacking production traces or when a domain expert is unavailable to manually annotate 50 traces for human-validated rubric alignment.