synthetic-workflow-polish-eval

Grades synthetic workflow polish outputs using a weighted rubric and outputs structured verdict YAML with calibration metadata.

1|2|Updated Apr 1, 2026
One-click install
npx skills add https://github.com/dimagi-internal/ace --skill synthetic-workflow-polish-eval
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: synthetic-workflow-polish-eval
Source: https://github.com/dimagi-internal/ace/tree/main/skills/synthetic-workflow-polish-eval
Command: npx skills add https://github.com/dimagi-internal/ace --skill synthetic-workflow-polish-eval

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill solves the problem of inconsistent, unstandardized grading for the synthetic workflow polish step, the most consequential stage for demo quality in ACE Phase 7 Plan B, which previously lacked strict validation that accounts for both text-based patch conformance and actual rendered visual output.

Core Features & Use Cases

  • 7-Dimension Weighted Rubric: Grades narrative-data coherence, patch quality, smoke-render success, domain language fit, mode honesty, and out-of-chain visual hierarchy and brand fit for comprehensive quality assessment.
  • Hard Blocker Enforcement: Automatically fails evals for critical issues like mismatched FLW names, broken smoke renders, and blocked visual judge verdicts to prevent low-quality demos from progressing.
  • Structured Verdict Output: Generates a standardized YAML verdict report with dimension scores, hard deduct triggers, and calibration metadata to support continuous rubric improvement.
  • Use Case: For ACE operators running Stage 4 of Phase 7 Plan B, this skill ensures polished synthetic workflows meet strict stakeholder-ready standards before demo deployment, eliminating subjective quality assessments.

Quick Start

Use the synthetic-workflow-polish-eval skill to grade the latest synthetic workflow polish run for the current Connect opportunity and generate a full structured verdict report.

Frequently Asked Questions about synthetic-workflow-polish-eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate synthetic workflow polish for demo readiness?

To evaluate synthetic workflow polish for demo readiness, grade step outputs against a standardized 7-dimension weighted rubric covering narrative coherence, patch quality, smoke render success, and visual fitness. This generates a structured verdict YAML with calibration metadata.

What is rubric scoring for visual judging and render validation?

Rubric scoring for visual judging evaluates rendered workflow screenshots against standardized criteria like visual hierarchy and brand fit. It uses a visual judge to capture out-of-chain visual fitness, ensuring polished workflows meet strict stakeholder-ready demo standards.

How do I enforce hard blockers for broken smoke renders and mismatched FLW names?

Enforce hard blockers for broken smoke renders and mismatched FLW names by automatically failing evaluations when critical issues are detected. This prevents low-quality demos from progressing through the synthetic workflow polish stage.

Can I generate a structured verdict report for patch conformance and domain language fit?

Yes, you can generate a structured YAML verdict report for patch conformance and domain language fit. This report includes dimension scores, hard deduct triggers, and calibration metadata to support continuous rubric improvement.

What are the limitations of automated polish evaluation for demo quality?

Automated polish evaluation relies on capturing rendered screenshots for visual judging and enforcing hard blockers for mismatched FLW names or blocked visual judge verdicts. It cannot assess subjective design aesthetics beyond the defined 7-dimension weighted rubric.

Does synthetic workflow polish evaluation work without dependencies?

Yes, synthetic workflow polish evaluation operates without external dependencies. It internally evaluates text-based patch conformance and uses a visual judge to score out-of-chain visual fitness from captured screenshots to produce a standardized verdict.