What problem does it solve?
Manually setting up LLM evaluators, scoring spans or experiment runs, and configuring continuous monitoring on Arize is time-consuming and prone to configuration errors like incorrect column mappings or missing credentials. This Skill eliminates that manual overhead by guiding you through end-to-end evaluation workflows with built-in validation and error handling.
Core Features & Use Cases
- LLM-as-Judge & Code Evaluator Management: Create, version, and update both LLM-powered template evaluators and deterministic code evaluators for tasks like hallucination detection, correctness scoring, and format validation.
- Project & Experiment Evaluation: Run one-time backfills or continuous scoring on live project spans, or score experiment dataset runs with custom column mappings to match your data schema.
- Use Case: For a RAG application deployed on Arize, use this Skill to automatically score retrieval relevance and answer correctness for every new user query, and backfill scores on historical traces to identify performance regressions.
Quick Start
Use the arize-evaluator skill to create a hallucination checker for your Arize project and run a backfill evaluation on the last 100 spans to validate scoring accuracy.