evaluation

Create scalable evaluation frameworks for agent systems with rubric-based scoring and test-set management.

Updated Jan 6, 2026
One-click install
npx skills add https://github.com/salmanparacha/speckitplus-calculator --skill evaluation-salmanparacha
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluation
Source: https://github.com/salmanparacha/speckitplus-calculator/tree/main/.claude/skills-nocontext/evaluation
Command: npx skills add https://github.com/salmanparacha/speckitplus-calculator --skill evaluation-salmanparacha

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

Evaluation of agent systems requires structured, repeatable measurement to distinguish quality from luck. This Skill provides a framework to design rubrics, test sets, and evaluation pipelines that quantify agent performance and guide improvements.

Core Features & Use Cases

  • Rubric-based scoring with per-dimension weights to capture multi-dimensional quality
  • Test-set management for organizing diverse evaluation scenarios and metrics
  • Evaluation runner that simulates or executes agent outputs and produces per-test results
  • Production monitoring to surface trends and alert when quality degrades in live use cases

Quick Start

Run a full evaluation cycle on a representative test set to generate actionable quality metrics.

Frequently Asked Questions about evaluation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate agent performance across non-deterministic runs?

Evaluating agent performance across non-deterministic runs requires a structured rubric-based scoring model with per-dimension weights. This framework simulates agent outputs and tracks improvements over time to deliver actionable quality insights.

What is rubric-based scoring for agent quality assurance?

Rubric-based scoring for agent quality assurance is a multi-dimensional evaluation method using per-dimension weights. It provides a structured, repeatable framework to measure agent performance and distinguish quality from luck.

How do I manage test sets for evaluating AI agent systems?

You manage test sets for evaluating AI agent systems by using test-set management features to organize diverse evaluation scenarios and metrics. This enables an evaluation runner to execute outputs and produce actionable per-test results.

Can I monitor production agent quality and alert on degradation?

Yes, you can monitor production agent quality and alert on degradation. The framework includes production monitoring to surface quality trends and alert you when live use cases experience performance degradation.

Does this agent evaluation framework support validating context engineering choices?

Yes, the agent evaluation framework supports validating context engineering choices. It applies a scalable evaluation pipeline with test-set management and rubric-based scoring to measure how context configurations perform.

What is the best way to track agent improvements over time?

The best way to track agent improvements over time is by running a full evaluation cycle on a representative test set. This generates actionable quality metrics using rubric-based scoring and production monitoring to surface trends.