evaluate-skill

Evaluate AI-generated outputs against predefined criteria for scoring and iteration decisions.

Updated Jan 20, 2026
One-click install
npx skills add https://github.com/maxoreric/sop-engine --skill evaluate-skill
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluate-skill
Source: https://github.com/maxoreric/sop-engine/tree/main/%24workflow.input.user_intent/skills/evaluate-skill
Command: npx skills add https://github.com/maxoreric/sop-engine --skill evaluate-skill

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) and scripts (resource) components.

What problem does it solve?

This Skill addresses the need for objective evaluation of AI-generated content, ensuring quality and identifying areas for improvement.

Core Features & Use Cases

  • Standardized Evaluation: Applies predefined criteria for consistent scoring.
  • Iterative Improvement: Determines if outputs meet standards or require further refinement.
  • Use Case: After an AI generates a report, use this Skill to evaluate its completeness, accuracy, and adherence to the original prompt, deciding whether to accept it or ask for revisions.

Quick Start

Use the evaluate skill to assess the attached report against the criteria in criteria.md.

Frequently Asked Questions about evaluate-skill

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate AI-generated outputs against predefined criteria?

To evaluate AI-generated outputs, you provide the output artifact and criteria definitions to assess quality and adherence, generating a score to decide whether to accept or request revisions.

What is the best way to assess code quality and report completeness for AI content?

Assessing code quality and report completeness involves applying standardized evaluation criteria to the AI-generated artifact, ensuring consistent scoring and identifying specific areas requiring iterative improvement.

How do I determine if AI-generated creative content meets specified standards?

Determining if creative content meets specified standards requires assessing the output artifact against your predefined criteria, yielding an evaluation score that guides iteration decisions for further refinement.

Do I need criteria definitions to score AI output?

Yes, you need criteria definitions to score AI output, as the evaluation mechanism requires both the criteria file and the output artifact to measure adherence and make iteration decisions.

Can I use standardized scoring to trigger iterative improvement of AI reports?

You can use standardized scoring to trigger iterative improvement by evaluating AI reports against predefined criteria, deciding whether the output meets standards or requires further refinement.