observability-engineer

Builds monitoring, logging, and tracing systems for production infrastructure.

1|Updated Jan 20, 2026
One-click install
npx skills add https://github.com/fakhriaditiarahman/Your-Skill-Agent --skill observability-engineer-fakhriaditiarahman
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability-engineer
Source: https://github.com/fakhriaditiarahman/Your-Skill-Agent/tree/main/.agent/skills/observability-engineer
Command: npx skills add https://github.com/fakhriaditiarahman/Your-Skill-Agent --skill observability-engineer-fakhriaditiarahman

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill addresses the critical need for robust monitoring, logging, and tracing systems to ensure the reliability, performance, and stability of production applications and infrastructure.

Core Features & Use Cases

  • Comprehensive Monitoring: Implement and manage metrics, alerts, and dashboards using tools like Prometheus, Grafana, and DataDog.
  • Distributed Tracing: Set up and analyze traces with Jaeger, Zipkin, and OpenTelemetry for deep insights into microservice interactions.
  • Log Management: Deploy and optimize ELK Stack, Loki, or Splunk for centralized logging and analysis.
  • Incident Response: Develop automated alerting, runbooks, and post-incident analysis workflows.
  • Use Case: You are launching a new microservices-based application and need to ensure you can quickly detect, diagnose, and resolve any performance issues or outages in production. This Skill will help you set up the necessary observability stack.

Quick Start

Use the observability-engineer skill to design a monitoring strategy for a new microservices architecture.

Frequently Asked Questions about observability-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I implement distributed tracing with OpenTelemetry?

Distributed tracing with OpenTelemetry is implemented by setting up tools like Jaeger and Zipkin to trace microservice interactions. This Skill guides you through configuring these tools to gain deep insights into your application's request flows.

What is the best way to centralize application logs using the ELK Stack?

Centralizing application logs using the ELK Stack provides a unified view of system events for analysis and troubleshooting. This Skill helps you deploy and optimize log management systems like ELK, Loki, or Splunk for your infrastructure.

Can I use Prometheus and Grafana for cloud-native observability?

Yes, Prometheus and Grafana can be used for cloud-native observability to collect metrics and visualize system performance. This Skill supports diverse technology stacks including cloud-native environments to proactively monitor infrastructure.

How do I create an incident response workflow with automated alerting?

Creating an incident response workflow with automated alerting involves developing runbooks and post-incident analysis procedures. This Skill builds these workflows to ensure you can quickly detect, diagnose, and resolve production outages.

When do I need centralized log management with Loki or Splunk?

Centralized log management with Loki or Splunk is needed when deploying microservices to aggregate logs for analysis and troubleshooting. This Skill enables you to deploy and optimize these stacks for comprehensive logging across enterprise applications.