What problem does it solve?
Reduces alert fatigue and missed incidents by automatically detecting deviations in service, infrastructure, and business metrics, correlating related signals across services, and producing prioritized remediation recommendations for operations teams.
Core Features & Use Cases
- Automated Baselines & Detection: Establish rolling baselines and detect anomalies using threshold, rate-of-change, and statistical rules.
- Correlation & Triage: Correlate metrics across services and infrastructure to produce a concise anomaly report with probable root causes and severity.
- Self-Improvement Loop: Measure alert-to-incident ratios, retune rules based on false positives, and update baselines and correlation patterns from postmortems.
- Use Case: Investigate a sudden latency spike by identifying affected services, correlating deployment and infra signals, and recommending targeted escalation and mitigation steps.
Quick Start
Run the anomaly-detector to scan recent service and infrastructure metrics, surface correlated anomalies, and generate a prioritized action and escalation recommendation.