crisai-eval-baseline-campaign

Evaluate LLM product quality across routing, grounding, fidelity, and cost metrics.

Updated Apr 18, 2026
One-click install
npx skills add https://github.com/crissdiamond/crisAI --skill crisai-eval-baseline-campaign
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: crisai-eval-baseline-campaign
Source: https://github.com/crissdiamond/crisAI/tree/main/.claude/skills/crisai-eval-baseline-campaign
Command: npx skills add https://github.com/crissdiamond/crisAI --skill crisai-eval-baseline-campaign

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill addresses the need for a comprehensive evaluation and acceptance baseline for crisAI's LLM product quality, encompassing routing accuracy, named-source resolution, source fit, grounding/citation, summary fidelity, artefact quality, policy gates, cost, latency, and human acceptance.

Core Features & Use Cases

  • Evaluation Framework: Provides a structured framework for evaluating LLM product quality across multiple dimensions.
  • Metrics Definition: Defines measurable metrics for each aspect of LLM performance.
  • Baseline Creation: Helps create a baseline for product quality based on defined metrics.
  • Use Case: For instance, when setting up an LLM product, this Skill can be used to ensure that the product meets the defined quality standards before release.

Quick Start

Load the skill with the command: uv run crisai run --skill crisai-eval-baseline-campaign

Frequently Asked Questions about crisai-eval-baseline-campaign

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I build an LLM evaluation baseline for routing accuracy and citation grounding?

LLM evaluation requires assessing routing accuracy, named-source resolution, source fit, grounding, summary fidelity, artefact quality, policy gates, cost, latency, and human acceptance. This Skill defines these measurable metrics to automate comprehensive product-quality assessment across all critical dimensions.

What metrics should I track for LLM product quality assurance?

You need a structured environment with predefined metrics and evaluation procedures. This Skill automates the evaluation process by providing a structured framework that defines measurable metrics across routing, grounding, and policy dimensions to create a product-quality baseline.

How do I set up an automated evaluation procedure for LLM summary fidelity?

This Skill uses an advanced implementation depth with scripts, references, and assets to automate LLM evaluation. It defines measurable metrics across multiple dimensions including routing, grounding, and latency, then runs structured procedures to establish a measurable product-quality baseline.

Does this LLM evaluation framework require any specific dependencies?

This LLM evaluation framework targets comprehensive product quality assessment across routing, grounding, artefact quality, cost, and human acceptance. It is designed for teams needing to ensure their LLM products meet defined quality standards before release.

Can I use this evaluation baseline to measure LLM cost and latency?

This Skill defines measurable metrics for LLM evaluation but requires a structured environment with predefined evaluation procedures. It focuses on comprehensive product-quality assessment across routing, grounding, and policy dimensions rather than isolated single-metric testing.