skill-judge

Evaluate installed AI skills across trigger accuracy, resolution, efficiency, and consistency.

Updated Jun 25, 2026
One-click install
npx skills add https://github.com/z1439527767/claude-config --skill skill-judge-z1439527767
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: skill-judge
Source: https://github.com/z1439527767/claude-config/tree/main/skills/imported/skill-judge
Command: npx skills add https://github.com/z1439527767/claude-config --skill skill-judge-z1439527767

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill solves the problem of keeping ineffective AI skills without evidence by systematically evaluating their real-world performance after repeated use.

Core Features & Use Cases

  • Multi-Axis Evaluation: Scores skills across trigger accuracy, problem resolution, token efficiency, side effects, and consistency.
  • Performance Verdicts: Produces weighted KEEP, IMPROVE, REPLACE, or REMOVE decisions based on measured outcomes.
  • Use Case: Analyze a frequently used coding assistant skill after several sessions to determine whether it should be promoted, refined, or removed.

Quick Start

Use the skill-judge skill to evaluate the effectiveness of a skill that has been used at least three times and generate a performance report.

Frequently Asked Questions about skill-judge

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate AI skill effectiveness after repeated usage?

To evaluate AI skill effectiveness, apply multi-axis scoring across trigger accuracy, problem resolution, token efficiency, side effects, and consistency to generate evidence-based performance verdicts.

When do I need to perform a skill quality audit for my AI workflows?

You need a skill quality audit when an installed AI skill has been used at least three times, requiring systematic evaluation of real-world performance to determine retention and improvement decisions.

How do I generate retention decisions like KEEP or REMOVE for installed AI skills?

Generate retention decisions by applying weighted scoring to measured outcomes from trigger analysis and regression tracking, producing structured KEEP, IMPROVE, REPLACE, or REMOVE verdicts for skill lifecycle management.

Can I measure token efficiency and trigger accuracy for a coding assistant skill?

Yes, you can measure token efficiency and trigger accuracy by analyzing a frequently used coding assistant skill after several sessions to determine whether it should be promoted, refined, or removed.

What is the best way to track performance regression in AI skill lifecycle management?

The best way to track performance regression is through systematic outcome review and efficiency measurement, applying structured evaluation criteria to compare consistency across multiple usage sessions.

What are the limitations of using weighted scoring for skill evaluation?

Weighted scoring for skill evaluation requires structured criteria and at least three usage instances to produce reliable verdicts, making it unsuitable for newly installed skills lacking sufficient historical evidence.