What problem does it solve?
This skill solves the problem of inconsistent, unstandardized grading for the synthetic workflow polish step, the most consequential stage for demo quality in ACE Phase 7 Plan B, which previously lacked strict validation that accounts for both text-based patch conformance and actual rendered visual output.
Core Features & Use Cases
- 7-Dimension Weighted Rubric: Grades narrative-data coherence, patch quality, smoke-render success, domain language fit, mode honesty, and out-of-chain visual hierarchy and brand fit for comprehensive quality assessment.
- Hard Blocker Enforcement: Automatically fails evals for critical issues like mismatched FLW names, broken smoke renders, and blocked visual judge verdicts to prevent low-quality demos from progressing.
- Structured Verdict Output: Generates a standardized YAML verdict report with dimension scores, hard deduct triggers, and calibration metadata to support continuous rubric improvement.
- Use Case: For ACE operators running Stage 4 of Phase 7 Plan B, this skill ensures polished synthetic workflows meet strict stakeholder-ready standards before demo deployment, eliminating subjective quality assessments.
Quick Start
Use the synthetic-workflow-polish-eval skill to grade the latest synthetic workflow polish run for the current Connect opportunity and generate a full structured verdict report.