What problem does it solve?
This Skill solves the time-consuming pain point of searching scattered internal and upstream documentation for TiKV raftstore and region lifecycle incident solutions, giving on-call engineering teams immediate access to pre-compiled diagnostic patterns, proven workarounds, and version-specific fix notes to resolve cluster stability and performance issues quickly.
Core Features & Use Cases
- Comprehensive Incident Coverage: Addresses common raftstore and region failure patterns including region split, add learner timeouts, snapshot apply delays, peer cleanup issues, large-region scheduling side effects, and raftstore-related panics or restarts.
- Actionable Diagnostic Guidance: Provides step-by-step initial checks, key metrics to monitor, diagnostic command examples, and proven workarounds for each known incident type.
- Deep-Dive Reference Materials: Includes detailed reference documents for complex failure modes like add-learner timeout loops and SST ingest latency spikes during data movement, plus links to related internal oncall cases and upstream TiKV GitHub issues.
Quick Start
Use the tikv-raftstore skill to troubleshoot your TiKV cluster's region split, snapshot apply delay, or raftstore panic incident by following the provided diagnostic steps and workarounds.