agent-evaluation

Evaluate LLM agent output quality using MLflow datasets and scorers.

Updated Feb 27, 2026
One-click install
npx skills add https://github.com/LaurentPRAT-DB/LPT_claude_config --skill agent-evaluation-laurentprat-db
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: agent-evaluation
Source: https://github.com/LaurentPRAT-DB/LPT_claude_config/tree/main/skills/agent-evaluation
Command: npx skills add https://github.com/LaurentPRAT-DB/LPT_claude_config --skill agent-evaluation-laurentprat-db

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill addresses the challenge of systematically evaluating and improving the output quality of Large Language Model (LLM) agents, ensuring they are accurate, cost-effective, and reliable.

Core Features & Use Cases

  • End-to-End Evaluation: Manages the complete workflow from setup to analysis.
  • MLflow Integration: Leverages MLflow's native APIs for datasets, scorers, and evaluation tracking.
  • Use Case: Improve an agent's tool selection accuracy by evaluating its performance on a curated dataset using predefined quality metrics and MLflow's tracing capabilities.

Quick Start

Use the agent-evaluation skill to evaluate the agent's performance on the provided dataset using the registered scorers.

Frequently Asked Questions about agent-evaluation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate LLM agent quality systematically?

To evaluate LLM agent quality systematically, this Skill manages the end-to-end workflow using MLflow's native APIs for datasets, scorers, and evaluation tracking to measure output accuracy and reliability.

How does MLflow work for evaluating LLM agent output?

MLflow evaluates LLM agent output by providing native APIs for datasets, scorers, and execution tracking. This Skill leverages these APIs to assess tool selection accuracy, answer quality, and cost reduction.

Can I use MLflow tracing to improve an agent's tool selection accuracy?

Yes, you can use MLflow tracing to improve an agent's tool selection accuracy by evaluating its performance on a curated dataset with predefined quality metrics to identify and correct incorrect responses.

What is the best way to reduce costs and incorrect responses in LLM agents?

The best way to reduce costs and incorrect responses in LLM agents is to apply systematic evaluation workflows using MLflow scorers and datasets, identifying specific output quality issues for targeted optimization.

Do I need a curated dataset to start evaluating agent performance with MLflow?

Yes, evaluating agent performance with MLflow requires a curated dataset. This Skill uses predefined quality metrics against your dataset to systematically assess tool selection accuracy and answer completeness.