What problem does it solve?
This Skill automates the comprehensive quality verification of Large Language Model (LLM) outputs across an entire pipeline, ensuring accuracy, relevance, and adherence to quality standards.
Core Features & Use Cases
- Automated Quality Gates: Implements scoring and validation for LLM-generated content at various pipeline phases.
- Evidence Score: Assigns a score (0-100) to assess the quality and grounding of LLM outputs, with defined actions for different score ranges (PASS, REVISE, REJECT).
- Multi-dimensional Quality Assessment: Evaluates outputs based on criteria like relevance, clarity, depth, bias, evidence, hallucination, and more.
- Use Case: Automatically verify the quality of job descriptions, candidate profiles, and interview questions generated by an LLM to ensure they meet predefined standards before proceeding to the next stage.
Quick Start
Use the quality-engine skill to verify the output quality of the LLM for phase P3 questions.