What problem does it solve?
This Skill detects evaluation-harness defects that can produce misleading zicato verdicts, including unwinnable ground truth, proxy-graded outputs, dead judges, inverted scoring, and nondeterministic results.
Core Features & Use Cases
- Known Baseline Validation: Compare trivially correct and trivially wrong stubs to verify that board expectations behave as intended.
- Harness Mechanics Audit: Check ground-truth winnability, graded-artifact fidelity, judge-fire counts, scalar arithmetic, ranking direction, determinism, and monotonicity scope.
- Telemetry-Based Diagnostics: Inspect loss profiles, generation scores, board declarations, event streams, and health findings to identify structural evaluation problems before tournaments or evolution.
- Use Case: After changing a zicato board, audit a deterministic baseline run to confirm that real agent outputs are graded, every declared judge fires, and better candidates receive the intended scalar ranking.
Quick Start
Use the zicato audit board skill to inspect the existing baseline run artifacts and determine whether the board can be trusted for tournament or evolution decisions.