What problem does it solve?
This Skill streamlines the creation of comprehensive evaluation plans for AI and LLM features, ensuring quality and enabling confident deployment.
Core Features & Use Cases
- Eval PRD Generation: Define clear requirements, scope, and acceptance thresholds for AI feature evaluations.
- Test Set & Taxonomy Creation: Develop structured golden test sets and detailed error taxonomies from observed failures.
- Rubric & Judge Planning: Design scoring rubrics and plan for human or LLM-as-judge execution.
- Use Case: You've developed a new AI chatbot. Use this Skill to design a complete evaluation plan, including a test set of user queries, a rubric for assessing response quality, and a plan for how to judge the responses, ensuring it meets safety and performance standards before launch.
Quick Start
Use the ai-evals skill to design evals for a customer-support reply drafting assistant, including a safety rubric and a golden set.