arize-evaluator

Create and run LLM-as-judge evaluators on Arize using the ax CLI.

5|6|Updated Jun 28, 2026
One-click install
npx skills add https://github.com/seldo/aiewf-2026-demo --skill arize-evaluator-seldo
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: arize-evaluator
Source: https://github.com/seldo/aiewf-2026-demo/tree/main/.agents/skills/arize-evaluator
Command: npx skills add https://github.com/seldo/aiewf-2026-demo --skill arize-evaluator-seldo

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires ax, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill simplifies the process of evaluating LLM outputs on Arize, providing tools for creating evaluators, running evaluations, managing tasks, and continuous monitoring.

Core Features & Use Cases

  • Evaluator Creation: Design and create evaluators for LLM-as-judge workflows.
  • Evaluation Execution: Run evaluations on spans or experiments.
  • Task Management: Manage tasks and trigger-run operations.
  • Column Mapping: Map columns for evaluators.
  • Continuous Monitoring: Monitor and improve evaluators continuously.
  • Use Case: Use this Skill to create an evaluator that checks the correctness of LLM responses in customer support interactions.

Quick Start

Use the arize-evaluator skill to create an evaluator for your project named 'customer-support-evaluator'.

Frequently Asked Questions about arize-evaluator

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I create LLM-as-judge evaluators on Arize?

You create LLM-as-judge evaluators on Arize by using this Skill to design, generate, and configure evaluators for your project, requiring an ax CLI and a configured Arize profile with an AI integration.

What do I need to run continuous monitoring for LLM evaluations?

To run continuous monitoring for LLM evaluations, you need the ax CLI installed and an Arize profile configured with an AI integration. This Skill automates the monitoring and management of your evaluators.

Can I run evaluations on spans and experiments using Arize AI integrations?

Yes, you can run evaluations on spans and experiments. This Skill automates the execution and management of LLM-as-judge evaluators on Arize to check the correctness of LLM responses.

How does column mapping work for LLM evaluators in Arize?

Column mapping for LLM evaluators maps your data columns to the required evaluator inputs. This Skill handles this mapping automatically as part of creating and managing your LLM-as-judge workflows on Arize.

Does the arize-evaluator skill work without the ax CLI?

No, the arize-evaluator skill requires the ax CLI to function. It also requires a configured Arize profile with an active AI integration to automate evaluator creation and evaluation execution.