evaluation-methodology

Define a methodology for measuring plugin and skill quality with static analysis, LLM judging, and Monte Carlo simulations.

1|Updated Apr 27, 2026
One-click install
npx skills add https://github.com/haxlys/skills --skill evaluation-methodology-haxlys
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluation-methodology
Source: https://github.com/haxlys/skills/tree/main/vendored/wshobson-agents/plugins/plugin-eval/skills/evaluation-methodology
Command: npx skills add https://github.com/haxlys/skills --skill evaluation-methodology-haxlys

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Provides a structured methodology to measure plugin and skill quality using multi-layer evaluation, enabling consistent assessment across products.

Core Features & Use Cases

  • Defines the three evaluation layers: static analysis, LLM judge, and Monte Carlo simulation.
  • Offers anchored rubrics for four judge dimensions and guidelines for score blending and badge decisions.
  • Enables calibration of qualitative signals and communication of results to partners and stakeholders.

Quick Start

Review the methodology to understand how PluginEval scores are calculated and how to interpret reports for your skill.

Frequently Asked Questions about evaluation-methodology

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate the quality of a plugin using a structured methodology?

To evaluate plugin quality, this methodology defines a multi-layer framework combining static analysis, LLM judging, and Monte Carlo simulations. It provides anchored rubrics across four dimensions to calculate scores, apply thresholds, and assign badges for consistent assessment.

What is the best way to calibrate an LLM judge for skill evaluation?

The best way to calibrate an LLM judge for skill evaluation is to apply the methodology's anchored rubrics across four defined dimensions. This process standardizes qualitative signals, ensuring accurate score blending and reliable badge decisions during assessment.

How are PluginEval scores calculated and interpreted for reports?

PluginEval scores are calculated by blending results from static analysis, LLM judge rubrics, and Monte Carlo simulations. Reports interpret these blended scores against predefined thresholds to determine badge criteria and communicate quality results to partners.

When do I need Monte Carlo simulations for measuring plugin quality?

You need Monte Carlo simulations for measuring plugin quality when assessing performance variability and robustness under randomized conditions. This methodology uses simulations as one of three core layers to ensure comprehensive scoring and reliable badge decisions.

Can I use these evaluation rubrics to communicate results to stakeholders?

Yes, you can use these evaluation rubrics to communicate results to stakeholders. The methodology includes specific guidelines for score blending, interpreting reports, and translating qualitative signals into clear badge criteria for partners.

What are the limitations of using static analysis for skill evaluation?

The limitation of using static analysis for skill evaluation is that it only covers one dimension of quality. This methodology requires combining static analysis with LLM judging and Monte Carlo simulations to achieve reliable scoring and comprehensive assessment.