eval-consistency

Evaluate persona roleplay consistency across dialogue scenarios with multi-dimension scoring.

107|21|Updated Apr 4, 2026
One-click install
npx skills add https://github.com/YIKUAIBANZI/forge-skill --skill eval-consistency
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: eval-consistency
Source: https://github.com/YIKUAIBANZI/forge-skill/tree/main/evals/eval-consistency
Command: npx skills add https://github.com/YIKUAIBANZI/forge-skill --skill eval-consistency

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps you evaluate whether a use-persona roleplay matches a given persona’s defining traits, including expression style, interaction behavior, and hard boundaries.

Core Features & Use Cases

  • Persona-driven scenario testing: Runs the same persona against multiple dialogue scenarios to generate replies.
  • Multi-dimension scoring: Scores each reply across five dimensions (length, catchphrases, punctuation/tone, interaction mode, and boundary compliance).
  • Actionable consistency report: Produces a per-scenario breakdown plus averages, pass/fail status (70+ target), common issues, and improvement suggestions.

Quick Start

Ask the AI to run the evaluation by saying: "/eval-consistency" with your persona and scenario inputs already available in the repository paths.

Frequently Asked Questions about eval-consistency

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I test persona roleplay consistency across multiple dialogue scenarios?

To test persona roleplay consistency, run the same persona against multiple dialogue scenarios to generate replies, self-score each across five dimensions out of 20, and output a structured consistency report with averages and recommendations.

What dimensions should a persona consistency scoring rubric evaluate?

A persona consistency scoring rubric should evaluate reply length, catchphrases, punctuation and tone, interaction mode, and boundary compliance, scoring each dimension out of 20 to determine pass or fail status.

Do I need specific input files to run persona consistency testing?

Yes, persona consistency testing requires loading persona.json chat-card layers (L0/L2/L4) and persona_consistency_cases.yaml inputs to generate replies in persona voice for scenario evaluation.

What is the passing score for a persona roleplay consistency benchmark report?

The passing score for a persona roleplay consistency benchmark report is 70 or above, which indicates that the persona reply generation meets the target consistency across the evaluated dialogue scenarios.

How does local-first LLM testing handle persona boundary compliance?

Local-first LLM testing handles persona boundary compliance by evaluating generated replies against the persona’s hard boundaries, scoring compliance out of 20, and flagging common issues in the final consistency report.

What is included in an actionable persona consistency report?

An actionable persona consistency report includes a per-scenario breakdown, dimension averages, pass or fail status against a 70-point target, common issues, and improvement suggestions for persona roleplay.