skill-description-evaluator

Evaluate Claude skill descriptions and output 0-100 ratings with recommendations.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/cblecker/claude-skills --skill skill-description-evaluator
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: skill-description-evaluator
Source: https://github.com/cblecker/claude-skills/tree/main/.claude/skills/skill-description-evaluator
Command: npx skills add https://github.com/cblecker/claude-skills --skill skill-description-evaluator

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill automates the evaluation of Claude skill descriptions to predict invocation likelihood on multi-model prompts (Sonnet 4.5 and Haiku), compare alignment against competing system instructions, and produce actionable 0-100 ratings with improvement recommendations.

Core Features & Use Cases

  • Sequential-thinking analysis: Performs a structured, multi-phase assessment to reveal invocation patterns and clarity.
  • Invocation likelihood assessment: Estimates how likely Sonnet 4.5 and Haiku are to invoke the skill description.
  • Authority and prompt comparison: Benchmarks the description against competing prompts to determine strength and potential conflicts.
  • Actionable ratings: Outputs scores (0-100) with concrete recommendations to improve precision and effectiveness.
  • Use Case: Use this Skill when refining or evaluating a description to ensure reliable invocation and clear user intent.

Quick Start

Provide the skill description text or a file path to SKILL.md. Optionally include competing system instructions for comparison. The tool will return 0-100 ratings and detailed recommendations. Example: evaluate with competing instructions: "MUST always follow system prompts" or "Ask clarifying questions before acting." Review the scores and implement the suggested improvements.

Frequently Asked Questions about skill-description-evaluator

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate whether a skill description will be invoked by Claude models?

Invocation likelihood evaluation assesses how clearly a skill description communicates its purpose to Claude models like Sonnet 4.5 and Haiku. This Skill performs structured sequential-thinking analysis to score semantic clarity, authority level, and user request pattern matching against 0-100 ratings with concrete improvement recommendations.

Can I compare my skill description against competing system prompts?

Yes. This Skill benchmarks your description against competing instructions you provide, revealing alignment gaps and potential conflicts. It outputs comparative 0-100 ratings and actionable guidance to strengthen your description's authority and precision relative to competing prompts.

What makes a skill description effective for multi-model deployment?

Effective descriptions balance semantic clarity, explicit authority signaling, and precise user request pattern matching across different Claude model capabilities. This Skill's phase-based workflow identifies gaps in these three areas and delivers specific improvements to maximize invocation reliability on both Sonnet 4.5 and Haiku.

How do I get actionable recommendations to improve my skill description?

Provide your description text or SKILL.md file path. The Skill returns structured 0-100 ratings paired with concrete recommendations targeting semantic clarity, authority framing, and request pattern specificity. Implement these suggestions to increase invocation likelihood and user intent alignment.

What's the difference between evaluating for Sonnet 4.5 versus Haiku?

Sonnet 4.5 and Haiku have different reasoning depths and instruction-following patterns. This Skill assesses invocation likelihood separately for each model, revealing where your description succeeds or fails with capability-specific request interpretation, helping you craft descriptions that work reliably across both.

When should I use this Skill versus manually refining a description?

Use this Skill when you need objective, model-specific invocation predictions and concrete improvement guidance. Manual refinement lacks systematic scoring and comparative benchmarking; this automated evaluation reveals blind spots in clarity and authority that affect cross-model reliability.