my-eval-plan

Design evaluation plans for AI/LLM features with scorers, datasets, and baselines.

4|1|Updated May 31, 2015
One-click install
npx skills add https://github.com/samcdavid/dotfiles --skill my-eval-plan
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: my-eval-plan
Source: https://github.com/samcdavid/dotfiles/tree/main/claude/skills/my-eval-plan
Command: npx skills add https://github.com/samcdavid/dotfiles --skill my-eval-plan

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Design evaluation plans for AI/LLM features before building them, ensuring structured criteria, guardrails, and measurable success across multiple platforms.

Core Features & Use Cases

  • Define evaluation dimensions (quality, safety, latency, robustness) and success criteria.
  • Design scorers (automated, LLM-judge, and human review) and calibration workflows.
  • Formulate dataset strategies (golden sets, edge cases, production samples, adversarial tests) and baseline targets.
  • Provide an end-to-end execution plan for offline and online evals.

Quick Start

Provide a feature description or ticket, and I will generate a complete eval plan including dimensions, scorers, datasets, baselines, and an execution plan.

Frequently Asked Questions about my-eval-plan

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design an LLM evaluation plan before deploying my AI feature?

To design an LLM evaluation plan, define evaluation dimensions like quality, safety, latency, and robustness, then establish measurable success criteria and baseline targets before building the actual evals.

What's the best way to structure pre-deployment datasets for AI evaluations?

The best way to structure pre-deployment datasets is to formulate a dataset strategy that includes golden sets, edge cases, production samples, and adversarial tests to ensure comprehensive coverage and robustness.

How do I set up LLM-judge scorers and calibration workflows for experiments?

Setting up LLM-judge scorers involves designing automated, LLM-judge, and human review scorers, alongside calibration workflows to consistently measure evaluation dimensions across your experiments.

Does this evaluation planning approach work with Braintrust or LangSmith?

Yes, this platform-agnostic approach works with Braintrust, LangSmith, custom harnesses, or manual reviews, providing an end-to-end execution plan for both offline and online pre-deployment evaluations.

Can I generate an execution plan for offline and online AI evals from a feature description?

Yes, you can generate a complete execution plan for offline and online AI evals by providing a feature description or ticket, which produces scorer definitions, dataset strategies, and baseline targets.