What problem does it solve?
This Skill eliminates the guesswork and lengthy investigation time for troubleshooting complex TiKV and TiCDC interaction failures, giving on-call engineers clear, curated guidance to resolve incidents faster.
Core Features & Use Cases
- Curated Known Failure Patterns: Access documented, real-world cases of common TiKV-TiCDC issues including CDC-triggered TiKV panics, changefeed lag after cluster scaling, and stale PD/store endpoint reuse, each with workarounds and related incident references.
- Targeted Diagnostic Workflow: Follow structured first checks to quickly differentiate between CDC metadata drift and TiKV runtime state root causes, reducing time spent on irrelevant troubleshooting steps.
- Deep Dive References: Access extended, detailed breakdowns of complex issues like TiKV address reuse-induced changefeed lag for more in-depth investigation when needed.
Use Case: For example, if your TiCDC changefeeds are experiencing increasing lag after a TiKV scale-out operation, this Skill helps you quickly identify if the root cause is stale store ID metadata and apply the correct restart workaround.
Quick Start
Use the tikv-cdc skill to investigate your TiKV-TiCDC interaction failure, identify the root cause, and apply the recommended workaround.