evaluation

Score completed deliverables against acceptance criteria with a 1–5 rubric.

1|Updated May 6, 2026
One-click install
npx skills add https://github.com/jacob-balslev/skill-graph --skill evaluation-jacob-balslev
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluation
Source: https://github.com/jacob-balslev/skill-graph/tree/main/marketplace/skills/evaluation
Command: npx skills add https://github.com/jacob-balslev/skill-graph --skill evaluation-jacob-balslev

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Evaluation prevents agents from prematurely declaring work complete by forcing an evidence-based check against the original request, acceptance criteria, and residual risks.

Core Features & Use Cases

  • Evidence-based completion scoring: Assign a disciplined 1–5 score using a visible score ceiling when required inputs or verification evidence are missing.
  • Coverage of completion gaps: Inventory request fit, completeness, correctness, verification evidence, residual risks, and reporting quality.
  • Evaluation-revision loop: Record findings with severity, evidence, required actions, and status, revise, re-verify, and rescore before accepting.
  • Boundary-aware routing: Use this for end-of-task completion evaluation, not for designing eval datasets/graders, debugging failures, diff line-by-line review, or pre-work methodology design.

Quick Start

Use the evaluation skill to score a completed deliverable against the original request and acceptance criteria using the available evidence packet, then list any required revisions and residual risks before marking it done.

Frequently Asked Questions about evaluation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I validate deliverable completion against acceptance criteria using evidence?

Deliverable completion validation requires collecting evidence and applying a 1–5 rubric with score ceilings to check request fit, correctness, and residual risks before accepting the work.

What is an evidence-based scorecard for reviewing agent task outcomes?

An evidence-based scorecard is a structured evaluation tool that inventories completion gaps, assigns severity to findings, and defines required actions to prevent prematurely declaring work done.

How do I run a revision and rescore loop for incomplete deliverables?

Run a revision and rescore loop by recording findings with required actions, revising the deliverable, re-verifying the evidence, and applying the 1–5 rubric again before final acceptance.

Can I use this completion review for documentation updates and skill upgrades?

Yes, completion review applies to documentation updates, skill upgrades, implementation deliverables, and other end-of-work artifacts that require a skeptical done gate before acceptance.

When should I not use a skeptical done gate for deliverable evaluation?

Deliverable evaluation is not intended for designing evaluation datasets, debugging failures, performing diff line-by-line review, or pre-work methodology design.

How do I handle missing verification evidence during a quality gate assessment?

During a quality gate assessment, apply a visible score ceiling to the 1–5 rubric when required inputs or verification evidence are missing, then list required actions to close the gap.