What problem does it solve?
This Skill eliminates the guesswork and lengthy investigation time for PD-related operational incidents, enabling on-call engineers to quickly identify root causes and apply proven workarounds to restore service.
Core Features & Use Cases
- Structured Triage Steps: Clear first checks to categorize incidents as ETCD, TLS, Dashboard, Prometheus scraping, or TSO-related, and distinguish between leader-only and cluster-wide degradation.
- Documented Known Cases: Detailed breakdowns of common failure patterns including ETCD capacity exhaustion, TLS rollout breakage, Dashboard unavailability, PD leader CPU spikes from TSO pressure, and Prometheus metric scrape gaps, each with verified workarounds and related incident tickets.
- Use Case Example: If your PD cluster loses service after ETCD backend usage exceeds its quota, this Skill provides the exact compaction, defragmentation, and alarm clearing steps to recover service without additional research.
Quick Start
Use the pd-operations skill to triage a PD leader CPU spike caused by TSO pressure and apply the recommended TSO follower proxy configuration change.