What problem does it solve?
This Skill eliminates the tedious, inconsistent manual work of grading OCS chatbot transcripts against standardized rubrics, ensuring consistent, auditable quality checks for Connect opportunities before launch and during ongoing operation.
Core Features & Use Cases
- Multi-Mode Grading: Supports three distinct modes: quick for 3-prompt Phase 5 shallow smoke tests, deep for pre-launch 5-dimension calibrated evaluation, and monitor for recurring Phase 9 performance trend tracking.
- Standardized Outputs: Produces machine-readable verdict YAML, human-readable evaluation reports, and gate briefs that integrate directly with upstream ACE orchestration gates and the opp-eval aggregation workflow.
- Calibrated Rubrics: Includes hard deduction rules, inflation guards, and auditable pre/post-cap scoring calibrated against per-opp ground truth to ensure consistent, defensible grading across runs.
- Use Case: A Connect team running Phase 5 QA can use quick mode to run a 3-prompt smoke gate in minutes to catch critical failures before advancing to deployment, while pre-launch teams use deep mode to validate the bot against 5 calibrated dimensions before releasing to Network Managers.
Quick Start
Use the ocs-chatbot-eval skill to grade the latest OCS chatbot transcript from your active Connect opportunity run in quick mode to validate it passes the Phase 5 shallow quality gate.