What problem does it solve?
This Skill turns a vague need to evaluate a GenAI application into a structured, implementation-ready evaluation plan with clear targets, datasets, metrics, and deployment guidance.
Core Features & Use Cases
- Evaluation Scope Discovery: Identifies agents, skills, and end-to-end flows that need testing before plans are written.
- Plan and Dataset Authoring: Produces organized markdown plans, test datasets, and model configuration files for repeatable evaluation.
- Platform Alignment: Centers the workflow on LangSmith for tracking and Azure Machine Learning with MLflow for scheduled runs and result visualization.
- Use Case: A team shipping an LLM feature can use this Skill to define what to test, how to score it, and how to operationalize regression evaluation across releases.
Quick Start
Ask this Skill to analyze your repository and produce a complete evaluation plan for your GenAI application, including datasets, metrics, model configuration, and Azure ML deployment guidance.