databricks-mlflow-evaluation

Evaluate MLflow 3 GenAI agents with automated workflows and Unity Catalog integration.

Updated May 31, 2026
One-click install
npx skills add https://github.com/thbeh/coding-agents-databricks-apps --skill databricks-mlflow-evaluation-thbeh
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: databricks-mlflow-evaluation
Source: https://github.com/thbeh/coding-agents-databricks-apps/tree/main/.claude/skills/databricks-mlflow-evaluation
Command: npx skills add https://github.com/thbeh/coding-agents-databricks-apps --skill databricks-mlflow-evaluation-thbeh

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires mlflow[databricks], openai, databricks-sdk, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill streamlines the evaluation process for MLflow 3 GenAI agents, reducing time spent on setup and troubleshooting common errors.

Core Features & Use Cases

  • Automated Evaluation Workflows: Follow pre-defined workflows for initial setup, trace analysis, and performance optimization.
  • End-to-End Monitoring: Set up trace ingestion, production monitoring, and continuous evaluation.
  • Custom Scorer Development: Guide to creating custom evaluation metrics for project-specific needs.
  • Unity Catalog Integration: Facilitates trace ingestion and production monitoring using Unity Catalog.
  • Judge Alignment: Aligns LLM judges with domain expert feedback for improved evaluation.
  • Automated Prompt Optimization: Enhances prompts using GEPA for automated prompt improvement.
  • Use Case: When you need to evaluate an MLflow GenAI agent and want to ensure it's performing accurately, efficiently, and safely.

Quick Start

Run the 'databricks-mlflow-evaluation' skill to evaluate an agent's performance on a dataset.

Frequently Asked Questions about databricks-mlflow-evaluation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate GenAI agents in MLflow 3?

To evaluate GenAI agents in MLflow 3, you can use automated workflows for initial setup, trace analysis, and performance optimization to ensure accurate and efficient agent operations.

How do I set up continuous monitoring for MLflow GenAI agents using Unity Catalog?

Continuous monitoring for MLflow GenAI agents is set up by integrating Unity Catalog to facilitate trace ingestion and production monitoring, enabling end-to-end evaluation of agent performance.

Can I create custom evaluation metrics for my MLflow GenAI agents?

Yes, you can create custom scorers to develop project-specific evaluation metrics, aligning LLM judges with domain expert feedback for improved evaluation accuracy.

What is the best way to optimize prompts for MLflow GenAI agents?

The best way to optimize prompts is using GEPA for automated prompt improvement, enhancing agent performance through automated workflows tailored to your specific evaluation needs.

Do I need specific libraries to run MLflow 3 GenAI agent evaluations?

Yes, evaluating MLflow 3 GenAI agents requires installing MLflow with Databricks, OpenAI, and the Databricks SDK to support the comprehensive evaluation framework and associated integrations.

Why does my MLflow GenAI agent evaluation require troubleshooting common errors?

MLflow GenAI agent evaluation often requires troubleshooting due to complex setup processes; this framework streamlines evaluation by reducing time spent on setup and resolving common configuration errors.