What problem does it solve?
Manually auditing TiCDC health requires selecting different metric rules for new and old runtime architectures, leading to inconsistent verdicts, missed critical issues, and longer incident response times for TiDB Cloud clusters.
Core Features & Use Cases
- Architecture-Aware Health Checks: Automatically classifies TiCDC as new or old architecture using metric probes, then applies the correct rule set for accurate verdicts.
- Comprehensive Evidence Collection: Aggregates metrics from Clinic, O11Y, and Prometheus data sources to generate per-changefeed and cluster-level health findings with clickable evidence links.
- Actionable Remediation Guidance: Provides prioritized risk lists and recommended next steps to resolve critical and warning issues quickly.
Use Case: A site reliability engineer can use this skill during a TiCDC incident to instantly get a health verdict, identify the root cause via linked metric queries, and follow the recommended remediation steps without manually reviewing dozens of metrics.
Quick Start
Use this skill to run a full TiCDC health inspection for your target cluster, receiving architecture classification, health verdicts, and recommended remediation actions.