What problem does it solve?
Centralizes and automates routine Datadog observability tasks so engineers can quickly query metrics, search logs, manage monitors and dashboards, post events, and schedule downtimes without manually calling disparate APIs.
Core Features & Use Cases
- Metric querying and exploration: Run time-series queries and list available metrics to diagnose performance regressions or track SLOs.
- Log search and analysis: Search logs with filters and pagination to investigate incidents or surface error trends.
- Monitor and dashboard management: Create, update, mute/unmute, and inspect monitors and dashboards to maintain alerting and visualization.
- Events, downtimes, hosts, and traces: Post events, schedule maintenance windows, list hosts, and retrieve traces for end-to-end incident workflows.
- Real-world example: Schedule a downtime for weekend maintenance, mute noisy monitors during the window, and create an event announcing the maintenance to stakeholders.
Quick Start
Use the datadog-automation skill to query the last hour error rate for service api and summarize whether it exceeded the alert threshold.