evaluation-methodology

Explain PluginEval skill quality scores and evaluation methodology.

Updated Apr 5, 2026
One-click install
npx skills add https://github.com/Jhabbig/Habbig --skill evaluation-methodology-jhabbig
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluation-methodology
Source: https://github.com/Jhabbig/Habbig/tree/main/.claude/plugins/wshobson/plugin-eval/skills/evaluation-methodology
Command: npx skills add https://github.com/Jhabbig/Habbig --skill evaluation-methodology-jhabbig

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill explains how PluginEval measures skill quality, helping you interpret scores, understand weak dimensions, and decide what to improve next.

Core Features & Use Cases

  • Scoring Interpretation: Break down composite scores, dimension grades, and badge thresholds into plain-language meaning.
  • Methodology Reference: Understand the three evaluation layers, anti-pattern penalties, Elo ranking, and blend weights behind the final result.
  • Improvement Guidance: Identify whether a skill needs better triggering, tighter scope, clearer output structure, or stronger references.
  • Use Case: A marketplace reviewer can use this Skill to explain why a submission earned Bronze instead of Silver and what changes would raise the score.

Quick Start

Use the evaluation-methodology skill to explain a PluginEval report, identify the weakest scoring dimensions, and recommend the most effective fixes.

Frequently Asked Questions about evaluation-methodology

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How does PluginEval scoring methodology work for skill quality assessment?

PluginEval scoring methodology evaluates skill quality across trigger accuracy, orchestration fitness, output quality, and scope calibration. It blends three evaluation layers using specific weights and applies Elo ranking to generate composite scores and badge thresholds.

What do the different PluginEval badge thresholds and dimension grades mean?

Badge thresholds and dimension grades in PluginEval reports reflect rubric anchors and anti-pattern penalties. This Skill breaks down composite scores into plain-language meaning, explaining why a submission earns a specific badge tier like Bronze or Silver.

How do I interpret PluginEval anti-pattern penalties and Elo ranking in my report?

Anti-pattern penalties reduce scores for poor triggering or scope calibration, while Elo ranking dynamically positions skills against each other. This Skill explains how these factors blend within the evaluation layers to determine your final quality score.

What is the best way to improve weak PluginEval dimensions and raise my score?

To improve weak PluginEval dimensions, review the methodology's rubric anchors and address anti-pattern penalties. This Skill identifies whether your submission needs better triggering, tighter scope, clearer output structure, or stronger references to increase its score.

Can I use this evaluation methodology to review trigger accuracy and orchestration fitness for any plugin?

Yes, the evaluation methodology applies to reviewing trigger accuracy, orchestration fitness, output quality, and scope calibration for plugin and skill assessments. It provides actionable improvement guidance based on the specific rubric anchors and blend weights evaluated.