What problem does it solve?
This Skill addresses the challenge of maintaining the health, performance, and reliability of cloud-based infrastructure and applications by providing comprehensive monitoring and observability.
Core Features & Use Cases
- Configures monitoring stacks: Sets up metrics, logs, and traces using tools like Prometheus, Grafana, CloudWatch, and OpenTelemetry.
- Establishes SLIs, SLOs, and SLAs: Defines and tracks key performance indicators and service level objectives.
- Implements alerting: Creates actionable alerts to minimize alert fatigue and ensure timely issue resolution.
- Use Case: A DevOps team can use this skill to set up a complete observability pipeline for their microservices, ensuring they can quickly detect and respond to performance degradations or outages.
Quick Start
Configure monitoring for our Kubernetes microservices on AWS using Prometheus and Grafana, focusing on API latency and error rates, and set up alerts to PagerDuty and Slack.