What problem does it solve? It resolves difficult, intermittent, or cross-layer software and infrastructure failures that resist routine fixes by proving the causal mechanism before applying the smallest durable repair. ## Core Features & Use Cases - Causal Investigation Workflow: Follows a structured state machine from intake through discovery, hypothesis testing, localization, proof, remediation, and verified reporting. - Evidence-Gated Completion: Maintains component verification matrices, incident timelines, and layered log ledgers so conclusions rest on evidence rather than correlation. - Bounded Remediation Budgets: Enforces attempt and time limits with anti-thrash gates requiring new evidence and new hypotheses before each retry. - Use Case: A Kubernetes workload fails only in production after a rollout. The skill preserves evidence, compares the deployment against vendor architecture, tests discriminating hypotheses, proves the earliest divergence, applies an owner-correct fix, and verifies it with the original reproducer. ## Quick Start Ask the agent to troubleshoot a persistent production failure by describing the expected behavior, the observed error, and the affected service or host.