skill-tester

Execute test scenarios for AI skills and generate detailed compliance reports.

244|75|Updated Dec 22, 2025
One-click install
npx skills add https://github.com/pavel-molyanov/molyanov-ai-dev --skill skill-tester-pavel-molyanov
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: skill-tester
Source: https://github.com/pavel-molyanov/molyanov-ai-dev/tree/main/skills/skill-tester
Command: npx skills add https://github.com/pavel-molyanov/molyanov-ai-dev --skill skill-tester-pavel-molyanov

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill automates the comprehensive testing of other AI skills, ensuring they function as intended and meet defined acceptance criteria. It provides detailed reports on performance and compliance.

Core Features & Use Cases

  • Automated Test Execution: Runs predefined scenarios against a target skill, comparing its output to a baseline.
  • Grader Agents: Utilizes specialized agents to analyze runner transcripts and evaluate acceptance criteria objectively.
  • Detailed Reporting: Generates structured reports highlighting skill value, issues, and ambiguities.
  • Use Case: Before deploying a new customer support skill, use the skill-tester to run through 50 common user queries, verifying its accuracy, helpfulness, and adherence to support protocols.

Quick Start

Run the skill tests for the 'customer-support' skill.

Frequently Asked Questions about skill-tester

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate skill validation and testing for AI agents?

Automated skill validation executes predefined test scenarios against a target AI skill, comparing its output to a baseline. It spawns parallel runners interacting as user personas to verify acceptance criteria and compliance.

What is quality assurance reporting for AI skills and how does it work?

Quality assurance reporting for AI skills utilizes grader agents to analyze runner transcripts and evaluate acceptance criteria objectively. It generates structured reports highlighting skill value, issues, and ambiguities.

How do I run test scenarios to evaluate agent performance before deployment?

To evaluate agent performance, you run predefined test scenarios against the target skill to compare its output with a baseline. The system spawns parallel runners and produces detailed reports on accuracy and compliance.

Can I use grader agents to check compliance with acceptance criteria?

Yes, grader agents analyze runner transcripts to objectively evaluate compliance with defined acceptance criteria. They identify skill value, issues, and ambiguities, providing structured evaluation reports.

What is the best way to test AI skills against a baseline output?

The best way to test AI skills against a baseline is by spawning parallel runners that interact as user personas. This automated testing compares skill-enabled runners against the baseline to grade compliance.

Does automated skill testing work for verifying customer support query accuracy?

Automated skill testing works for verifying customer support query accuracy by running common user queries against the skill. It validates accuracy, helpfulness, and adherence to support protocols before deployment.