What problem does it solve?
This skill addresses the instability of CI performance and correctness thresholds by providing a statistically sound, automated calibration process that distinguishes between genuine model regressions and transient host-level noise.
Core Features & Use Cases
- Destructive Round Rejection: Automatically identifies and discards rounds contaminated by host contention or cold caches using robust statistical tests (MAD and gap analysis).
- Strict Provenance & Readiness: Ensures every calibration run is tied to a specific git commit and environment fingerprint, preventing the use of stale or incompatible data.
- Use Case: When a new model version causes CI performance tests to flake, use this skill to generate a reliable, worst-of-N threshold report that accounts for hardware variability, ensuring your CI assertions remain stable and meaningful.
Quick Start
Use the tune-ci-thresholds skill to run a full calibration for the omni model on the assigned GPU group and generate a report.