evaluation-methodology

Explain PluginEval's three-layer scoring, badge thresholds, and anti-pattern flags.

Updated Apr 19, 2026
One-click install
npx skills add https://github.com/ArogyaReddy/https-github.com-wshobson-agents --skill evaluation-methodology-arogyareddy
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluation-methodology
Source: https://github.com/ArogyaReddy/https-github.com-wshobson-agents/tree/main/plugins/plugin-eval/skills/evaluation-methodology
Command: npx skills add https://github.com/ArogyaReddy/https-github.com-wshobson-agents --skill evaluation-methodology-arogyareddy

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

PluginEval methodology clarifies how plugin and skill quality are measured, enabling objective assessment and improvement.

Core Features & Use Cases

  • Three-layer evaluation framework covering static analysis, LLM judge, and Monte Carlo simulations to produce a robust quality score.
  • Anchored rubrics and badge thresholds that translate scores into actionable quality badges for partners and marketplaces.
  • Guidance for improvement including anti-patterns and practical tips to raise triggering accuracy, orchestration fitness, and output quality.

Quick Start

Review the frontmatter and rubrics to understand PluginEval's quality measurement framework.

Frequently Asked Questions about evaluation-methodology

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How does plugin quality evaluation work in PluginEval?

Plugin quality evaluation in PluginEval uses a three-layer framework: static analysis, LLM judge, and Monte Carlo simulations. These layers combine to produce a robust composite score measuring overall plugin and skill quality across ten dimensions.

How do I calibrate scoring thresholds for a marketplace?

To calibrate scoring thresholds for a marketplace, apply the anchored rubrics defined in the PluginEval methodology. These rubrics translate raw composite scores into actionable quality badges, ensuring objective and consistent partner evaluations.

What are anti-pattern flags in plugin evaluation?

Anti-pattern flags in plugin evaluation identify specific design flaws affecting triggering accuracy, orchestration fitness, and output quality. Recognizing these flags provides practical guidance to improve plugins and raise overall quality scores.

How do I explain quality badges to external partners?

To explain quality badges to external partners, reference the anchored rubrics and badge thresholds from the PluginEval methodology. Badges represent a plugin's composite score, translating evaluation results into clear, actionable quality indicators.

What scoring dimensions are used to measure plugin quality?

Plugin quality is measured across ten distinct scoring dimensions defined within the PluginEval methodology. These dimensions evaluate triggering accuracy, orchestration fitness, and output quality to calculate a comprehensive composite score.

Can I use PluginEval methodology for skill quality assessment?

Yes, you can use the PluginEval methodology for skill quality assessment. The framework measures both plugins and skills through its three-layer evaluation process, applying the same ten scoring dimensions and composite formula.