observability-engineer

Design observability signals, dashboards, and alerting for production services.

1|Updated Sep 11, 2025
One-click install
npx skills add https://github.com/Dhumitech/DHUMI-AI-RESOURCE --skill observability-engineer-dhumitech
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability-engineer
Source: https://github.com/Dhumitech/DHUMI-AI-RESOURCE/tree/main/AI-Engineer-planner-Skills/07-deploy/observability-engineer
Command: npx skills add https://github.com/Dhumitech/DHUMI-AI-RESOURCE --skill observability-engineer-dhumitech

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Production systems require reliable visibility. This Skill helps design and implement comprehensive observability strategies to monitor, trace, log, and alert for enterprise-scale applications.

Core Features & Use Cases

  • Design monitoring and tracing architectures that provide end-to-end visibility across services.
  • Implement dashboards, alerts, and runbooks aligned to SLOs for incident response.
  • Real-world scenario: Deploy a starter observability stack for a multi-service web application to reduce MTTR by correlating traces, metrics, and logs.

Quick Start

Define your production services, instrument core signals (metrics, logs, traces), and implement starter dashboards and alerting aligned to SLOs.

Frequently Asked Questions about observability-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I establish production observability for enterprise microservice architectures?

To establish production observability for enterprise microservice architectures, you design comprehensive signals, instrumentation, and dashboards. This provides end-to-end visibility across services to improve reliability and reduce mean time to resolution.

What is the best way to implement OpenTelemetry instrumentation across containerized workloads?

The best way to implement OpenTelemetry instrumentation across containerized workloads involves defining core signals, instrumenting metrics, logs, and traces, and implementing starter dashboards. This secures monitoring data and ensures reliable visibility.

Can I use this approach to design SLI and SLO tracking for serverless environments?

Yes, you can use this approach to design SLI and SLO tracking for serverless environments. It enables the definition and tracking of service level indicators, aligning dashboards and alerting to your reliability objectives.

How do I set up alerting and runbook automation for incident response?

Setting up alerting and runbook automation for incident response requires implementing alerts aligned to SLOs. This approach correlates traces, metrics, and logs to reduce MTTR and automate operational procedures during production incidents.

When do I need to correlate traces, metrics, and logs for a multi-service web application?

You need to correlate traces, metrics, and logs for a multi-service web application when deploying a starter observability stack. This correlation provides comprehensive visibility, reduces MTTR, and ensures reliable monitoring across all services.

Does comprehensive observability work without secure handling of monitoring data?

No, comprehensive observability requires secure handling of monitoring data to satisfy functional requirements. Designing signals and instrumentation across enterprise-scale applications must include secure data practices to maintain reliability.