evaluation

Evaluate agent systems with multidimensional rubrics and automated or human modes.

Updated Nov 16, 2025
One-click install
npx skills add https://github.com/mhintz1980/ptl-lova --skill evaluation-mhintz1980
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluation
Source: https://github.com/mhintz1980/ptl-lova/tree/main/docs/agent-skills/skills/evaluation
Command: npx skills add https://github.com/mhintz1980/ptl-lova --skill evaluation-mhintz1980

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

Evaluation of agent systems requires formal frameworks to measure performance, compare configurations, and drive iterative improvement across cycles.

Core Features & Use Cases

  • Multidimensional Rubrics: Balanced scoring across fidelity, completeness, citation quality, source credibility, and tool efficiency.
  • Test Set Management: Create, filter, and analyze evaluation test sets with complexity distributions and tagging.
  • Production Monitoring: Lightweight sampling and dashboards that surface trends and alert on quality degradation.

Quick Start

Run a standardized evaluation suite to measure agent performance across multiple tasks.

Frequently Asked Questions about evaluation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate agent systems to measure performance improvements?

You can evaluate agent systems by applying multi-dimensional rubrics that score fidelity, completeness, citation quality, and tool efficiency. This framework tracks trends over time to measure performance and drive iterative improvement across cycles.

What metrics should I use for LLM evaluation of agent systems?

LLM evaluation metrics should cover fidelity, completeness, citation quality, source credibility, and tool efficiency. Using balanced multi-dimensional rubrics ensures comprehensive performance measurement for agent systems across various tasks.

How do I create and manage test sets for evaluating agent performance?

You can create, filter, and analyze evaluation test sets by applying complexity distributions and tagging. This test set management enables systematic testing and context engineering validation across your production pipelines.

Can I monitor production agent quality using automated and human evaluation modes?

Yes, production monitoring supports both automated and human evaluation modes. It uses lightweight sampling and dashboards to surface trends and alert on quality degradation in your agent systems.

What's the best way to set up continuous quality gates for production pipelines?

The best way to establish continuous quality gates is by running standardized evaluation suites with multi-dimensional rubrics. This approach enables systematic testing and tracks performance trends to prevent degradation across production pipelines.

When do I need formal evaluation frameworks for agent systems?

You need formal evaluation frameworks when you must compare configurations, validate context engineering, and drive iterative improvement. They provide structured quality assurance to measure performance accurately across multiple cycles.