What problem does it solve?
This Skill solves the problem of test prompt sets that are structurally valid but fail to stress-test real-world LLO chatbot edge cases, leading to eval results that give false confidence in bot performance. It ensures prompt suites are not just faithful to thin PDDs, but actually cover safety-critical, out-of-knowledge-base, and adversarial scenarios a real LLO supervisor would encounter.
Core Features & Use Cases
- 7-Dimension Quality Grading: Evaluates test prompts across expected-answer specificity, adversarial-prompt quality, archetype coverage, prompt phrasing realism, expected-tag correctness, escalation-prompt quality, and out-of-chain failure-mode coverage.
- Out-of-Chain Fitness Checks: Enforces a hard floor for failure-mode coverage to prevent thin, PDD-aligned prompt suites from passing evals without testing real-world edge cases.
- Structured Verdict Generation: Produces a standardized YAML verdict with weighted scores, auto-surfaced blockers/warnings, and calibration tracking for consistent eval results.
- Use Case: Eval teams building LLO chatbot test suites can use this Skill to verify their prompts will catch real bot failures before deployment, avoiding costly post-launch issues.
Quick Start
Use the pdd-to-test-prompts-eval skill to grade the quality of your test prompt set for the current Connect opportunity and generate a structured eval verdict.