What problem does it solve?
This Skill helps you investigate production incidents by correlating logs, traces, metrics, and code to find root cause and recommend mitigation. It is useful for debugging live issues, performing root cause analysis, and handling production anomalies.
Core Features & Use Cases
- Correlation of Observability Data: Correlate logs, traces, metrics, and code to identify the root cause of production issues.
- Connector Readiness: Load available observability connectors and verify their setup.
- Triage: Extract symptoms, services, time windows, identifiers, and suspicions from user input.
- Strategy: Decide on data sources and query order to investigate the issue.
- Gather Signals: Execute queries for logs, APM, traces, codebase, and deploy history.
- Correlate: Build a concrete evidence chain to map user-facing IDs to internal IDs and compare before/after deploy versions.
- Mitigate First: Advise mitigation steps before full root cause analysis.
- RCA Report: Generate a structured report detailing contributing factors and actions taken.
Quick Start
Load the investigate skill and provide the symptom, service, time window, identifiers, and suspicion details for the incident you are investigating.