evals-router

Route LLM evaluation-method requests to matching local workflow files.

8|4|Updated Jan 15, 2026
One-click install
npx skills add https://github.com/jscraik/Agent-Skills --skill evals-router
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evals-router
Source: https://github.com/jscraik/Agent-Skills/tree/main/utilities/evals-router
Command: npx skills add https://github.com/jscraik/Agent-Skills --skill evals-router

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) and workflows (resource) components.

What problem does it solve?

This Skill streamlines the process of directing complex LLM evaluation tasks to the most appropriate specialized workflow, ensuring efficient and accurate problem-solving.

Core Features & Use Cases

  • Task Routing: Intelligently directs requests for LLM evaluation design, auditing, debugging, scaling, error analysis, judge prompt design, evaluator validation, RAG evaluation, synthetic data generation, or human review interfaces.
  • Efficiency: Prevents misallocation of resources by identifying the narrowest, most trustworthy workflow for immediate action.
  • Use Case: When a user asks to "audit our current eval setup and tell me what is missing before we trust the scores," this Skill routes them to the eval-audit workflow.

Quick Start

Use the evals-router skill to audit our LLM eval pipeline and tell us what is missing before we trust the scores.

Frequently Asked Questions about evals-router

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I route LLM evaluation tasks to the correct workflow?

Route LLM evaluation tasks by matching user goals and available artifacts to specific local workflow files, providing prerequisite and trust-boundary guidance for accurate execution.

What is the best way to audit an LLM evaluation setup before trusting scores?

Audit an LLM evaluation setup by directing requests to the eval-audit workflow, which identifies missing components and validates trust boundaries before relying on evaluation scores.

Can I use this for RAG evaluation and synthetic eval data generation?

RAG evaluation and synthetic eval data generation are supported workflows, triggered by matching your specific evaluation goals and available artifacts to the corresponding local workflow files.

How does LLM evaluator validation and judge prompt design work?

LLM evaluator validation and judge prompt design are handled by routing your specific request to dedicated workflows, ensuring explicit prerequisite checks and trust-boundary guidance are applied.

Do I need existing artifacts to validate my LLM evaluation pipeline?

Existing artifacts are required to validate an LLM evaluation pipeline, as the routing mechanism matches your available artifacts to the narrowest, most trustworthy workflow for immediate action.

What are the limitations of using automated routing for LLM evaluation?

Automated routing for LLM evaluation relies on matching user goals and artifacts to predefined workflows, meaning it cannot process tasks outside its supported scope of eval design, auditing, and review tooling.