What problem does it solve?
Manually setting up and managing human annotation workflows for Arize LLM observability and model evaluation projects is slow, inconsistent, and prone to configuration errors that break labeling pipelines and delay model performance insights.
Core Features & Use Cases
- Annotation Schema Management: Create, update, and delete categorical, continuous, and freeform label schemas to standardize human feedback across your Arize projects.
- Review Queue Configuration: Set up and manage human review annotation queues with custom instructions, reviewer assignments, and linked label schemas to streamline labeling workflows.
- Bulk Span Annotation: Apply human annotations to large sets of Arize project spans via the Python SDK or CLI to speed up model evaluation and trace monitoring.
- Use Case: For example, if you are evaluating the performance of a new customer support LLM, use this skill to quickly create a correctness label schema, assign a review queue to your support team, and bulk apply reviewer labels to all conversation spans for performance analysis.
Quick Start
Use the arize-annotation skill to create a categorical correctness label schema, set up a review queue for your team, and bulk apply reviewer labels to all spans in your current Arize project.