skill-tester

Create and apply evaluation rubrics for AI skills with test scenarios.

27|5|Updated Feb 24, 2026
One-click install
npx skills add https://github.com/harvard-lil/skills-hub-demo --skill skill-tester-harvard-lil
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: skill-tester
Source: https://github.com/harvard-lil/skills-hub-demo/tree/main/skills/skill-developer/skill-tester
Command: npx skills add https://github.com/harvard-lil/skills-hub-demo --skill skill-tester-harvard-lil

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill helps developers ensure their AI skills are effective, reliable, and meet defined quality standards by creating and applying evaluation rubrics.

Core Features & Use Cases

  • Rubric Creation: Guides users through defining structural criteria, pedagogical goals, and anti-patterns for evaluating AI skills.
  • Test Scenario Design: Helps craft specific user interactions to test skill performance in various situations.
  • Conversation Evaluation: Assesses actual skill usage against a defined rubric, providing detailed feedback.
  • Use Case: A skill developer wants to ensure their new 'legal-research-assistant' skill consistently provides accurate citations and avoids giving legal advice. They use this Skill to create a rubric with specific checks for citation format and a clear anti-pattern for providing advice.

Quick Start

Use the skill-tester to create a rubric for the 'socratic-tutor' skill.

Frequently Asked Questions about skill-tester

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I create an evaluation rubric for an AI skill?

To create an evaluation rubric, define structural criteria, pedagogical goals, and anti-patterns for the AI skill. This establishes clear quality standards and specific checks to measure skill performance against defined objectives.

What is skill testing and when do I need to evaluate AI performance?

Skill testing is the process of assessing AI performance against defined quality standards using rubrics. You need it to ensure your AI skills are effective, reliable, and consistently meet pedagogical objectives before deployment.

How do I design test scenarios to evaluate AI skill performance?

Design test scenarios by crafting specific user interactions that probe various situations. This helps assess actual skill usage against the defined rubric and provides detailed feedback on performance.

Can I evaluate both automated and human interactions against a rubric?

Yes, you can support both automated and human evaluation of skill performance. The rubric defines quality criteria and anti-patterns, enabling comprehensive conversation evaluation against consistent standards.

What do I need to know to generate comprehensive rubrics for AI skills?

Generating comprehensive rubrics requires an understanding of skill mechanics and pedagogical objectives. This knowledge allows you to define accurate structural criteria, anti-patterns, and specific test scenarios for evaluation.

What is the best way to prevent an AI skill from exhibiting anti-patterns?

The best way to prevent anti-patterns is to explicitly define them within the evaluation rubric. This allows you to assess actual skill usage and ensure the AI consistently avoids undesirable behaviors like giving unauthorized advice.