evaluation

Evaluate context-engineered agents using configurable weighted rubrics and test sets.

1|Updated Jan 27, 2026
One-click install
npx skills add https://github.com/phonowell/mimikit --skill evaluation-phonowell
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluation
Source: https://github.com/phonowell/mimikit/tree/main/.agents/skills/context-engineering-collection/skills/evaluation
Command: npx skills add https://github.com/phonowell/mimikit --skill evaluation-phonowell

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

Evaluation frameworks for context-engineered agents that enable repeatable, multi-dimensional assessments of agent outputs, improving reliability and insight across variations in context and tooling.

Core Features & Use Cases

  • A configurable rubric with weighted dimensions (factual accuracy, completeness, citation accuracy, source quality, tool efficiency).
  • Test-set management and an evaluation runner to compare different agent configurations over time.
  • Production monitoring hooks to track quality signals on live interactions and trigger alerts.

Quick Start

Install and instantiate the evaluation pipeline, then run a standard test set through the runner to generate a summary.

Frequently Asked Questions about evaluation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate agent performance using multi-dimensional rubrics?

You evaluate agent performance by applying a structured framework with configurable rubrics, using weighted dimensions like factual accuracy and tool efficiency to quantify outputs across development, testing, and production.

What is the best way to monitor agent quality signals in production?

You monitor agent quality in production by attaching real-time monitoring hooks to live interactions, which track quality signals and trigger alerts based on your configured evaluation rubrics.

How do I compare different agent configurations over time?

You compare agent configurations over time by using test-set management and an evaluation runner to process standard test sets, generating summaries that highlight performance variations across context and tooling.

What dimensions should I include in an agent evaluation rubric?

An agent evaluation rubric should include weighted dimensions like factual accuracy, completeness, citation accuracy, source quality, and tool efficiency to ensure repeatable, multi-dimensional assessments of agent outputs.

Can I use this evaluation framework for context-engineered agents?

Yes, this evaluation framework is specifically designed for context-engineered agents, enabling repeatable, multi-dimensional assessments that improve reliability and insight across variations in context and tooling.