What problem does it solve?
Provides structured, proactive system analysis to surface reliability, performance, and architectural risks that are often missed by reactive incident response. It helps teams discover recurring patterns across incidents, prioritize technical risk, and decide when a finding requires a design-level escalation versus a point fix.
Core Features & Use Cases
- Discovery-driven System Mapping: Scans code, docs, and artifacts to build or refresh a system map before any deep analysis.
- Multi-mode Analysis: Run health assessments, pattern recognition across incidents, component deep dives, risk scoring, and architecture reviews with interactive checkpoints.
- Design Escalation & Reporting: Detects when issues need architectural redesign, offers inline trade-offs, and recommends running a deeper /design workflow; produces structured Markdown reports for tracking and handoff.
- Use Case: Run a patterns analysis on the last 30 days of incidents to identify a recurring timeout hotspot, escalate it for design trade-offs, and generate a prioritized findings report for the engineering team.
Quick Start
Run a full system health assessment focusing on reliability and observability and save the results to docs/analysis.