ywc-toolkit-eval

Evaluate Claude Code skills and agents with scoring and prioritized backlogs.

8|1|Updated May 13, 2026
One-click install
npx skills add https://github.com/yongwoon/ywc-agent-toolkit --skill ywc-toolkit-eval
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ywc-toolkit-eval
Source: https://github.com/yongwoon/ywc-agent-toolkit/tree/main/.claude/skills/ywc-toolkit-eval
Command: npx skills add https://github.com/yongwoon/ywc-agent-toolkit --skill ywc-toolkit-eval

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill addresses the challenge of maintaining high-quality standards across a large collection of Claude Code skills and agents by automating the evaluation, scoring, and improvement cycle.

Core Features & Use Cases

  • Graded Scorecarding: Evaluates skills and agents across six quality axes, including activation accuracy, structural compliance, and behavioral efficacy.
  • Prioritized Backlog: Automatically ranks the weakest items in your toolkit to provide a clear, actionable improvement path.
  • Use Case: Run this Skill after a major update to your toolkit to identify which skills have become redundant or have drifted from their intended behavioral contract, ensuring your catalog remains lean and accurate.

Quick Start

Use the ywc-toolkit-eval skill to perform a full quality assessment of all toolkit skills and agents and generate a prioritized improvement backlog.

Frequently Asked Questions about ywc-toolkit-eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate the quality of Claude Code skills and agents?

To evaluate Claude Code skills and agents, use a two-tier scoring harness that performs structural compliance checks and behavioral validation. This graded scorecard approach assesses activation precision across six quality axes to maintain catalog integrity.

What is the best way to identify redundant Claude Code skills in my toolkit?

The best way to identify redundant skills is by running an evaluation that generates a prioritized backlog. It automatically ranks the weakest items and tracks historical performance trends to highlight drifted behavioral contracts.

How does automated agent scoring drive continuous improvement for Claude Code?

Automated agent scoring drives continuous improvement by applying mechanical and judgment-based evaluations to custom agents. It identifies behavioral efficacy issues and operational drift, producing actionable feedback to optimize toolkit performance.

When should I perform a quality assessment on my custom agents and skills?

You should perform a quality assessment after major toolkit updates. This detects structural compliance issues and activation accuracy drift, ensuring your catalog remains lean, accurate, and free of redundant skills or agents.

Can I minimize token waste by grading my Claude Code toolkit?

Yes, grading your Claude Code toolkit minimizes token waste by identifying and removing redundant skills. The evaluation process ensures operational efficacy and structural compliance, keeping your catalog lean and efficient.