evaluate

Evaluate AI pipeline outputs using a weighted rubric and log iteration results.

Updated Jan 31, 2026
One-click install
npx skills add https://github.com/ctoscano/cw-hackathon-3 --skill evaluate-ctoscano
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluate
Source: https://github.com/ctoscano/cw-hackathon-3/tree/main/.claude/skills/evaluate
Command: npx skills add https://github.com/ctoscano/cw-hackathon-3 --skill evaluate-ctoscano

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill helps you automatically audit AI pipeline outputs and propose targeted, iterative improvements to prompts and configurations, enabling more consistent quality across pipelines such as DAP notes and intake questionnaires.

Core Features & Use Cases

  • Automated evaluation of pipeline outputs using a weighted rubric and expert personas.
  • Generate concrete improvement suggestions and track iteration history across runs.
  • Supports multiple pipelines by loading pipeline references and identifying latest outputs.

Quick Start

Run a full evaluation for the DAP notes pipeline with /evaluate run dap. You can also compare versions with /evaluate review dap.

Frequently Asked Questions about evaluate

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automatically evaluate AI pipeline outputs for quality?

You can evaluate AI pipeline outputs automatically by applying a weighted rubric to the generated results. This process audits outputs from pipelines like DAP notes and intake questionnaires to propose targeted iterative improvements.

What is the best way to iteratively improve AI prompts and configurations?

The best way to iteratively improve AI prompts is to log evaluation results from a weighted rubric across runs. Tracking this iteration history guides subsequent edits to your configurations for consistent quality.

Can I use a weighted rubric to audit DAP notes and intake questionnaire pipelines?

Yes, you can use a weighted rubric to audit DAP notes and intake questionnaire pipelines. The evaluation process loads pipeline references, identifies latest outputs, and applies expert personas to assess quality.

How do I run a full evaluation for the DAP notes pipeline?

To run a full evaluation for the DAP notes pipeline, execute the run command. This triggers the automated evaluation process, which reads references and logs iteration results to guide prompt edits.

How do I compare versions when evaluating AI pipeline outputs?

To compare versions when evaluating AI pipeline outputs, use the review command. This allows you to review previous iteration logs and evaluation results to track improvements across pipeline runs.

Does the evaluation process support multiple AI pipelines?

Yes, the evaluation process supports multiple AI pipelines by loading pipeline references and identifying latest outputs. It applies across DAP notes, intake questionnaires, and any future pipelines using the rubric.