observability-engineer

Build monitoring, logging, and tracing systems with Prometheus, Grafana, and OpenTelemetry.

23|2|Updated Jan 19, 2026
One-click install
npx skills add https://github.com/herdiansah/Antigravity-Skills-Master --skill observability-engineer-herdiansah
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability-engineer
Source: https://github.com/herdiansah/Antigravity-Skills-Master/tree/main/.agent/skills/observability-engineer
Command: npx skills add https://github.com/herdiansah/Antigravity-Skills-Master --skill observability-engineer-herdiansah

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill addresses the challenge of building robust, scalable, and cost-effective monitoring, logging, and tracing systems essential for maintaining the reliability and performance of modern applications.

Core Features & Use Cases

  • Comprehensive Monitoring: Implement solutions using Prometheus, Grafana, DataDog, and cloud-native tools.
  • Distributed Tracing: Set up Jaeger, Zipkin, and OpenTelemetry for deep application insight.
  • Log Management: Deploy ELK Stack, Loki, or Splunk for centralized log aggregation and analysis.
  • Incident Response: Design alerting, PagerDuty integration, and automated runbooks.
  • Use Case: You need to establish a complete observability stack for a new microservices-based e-commerce platform to proactively identify and resolve performance bottlenecks before they impact customers.

Quick Start

Design a comprehensive monitoring strategy for a microservices architecture with 50+ services.

Frequently Asked Questions about observability-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I build a comprehensive observability stack for a microservices architecture?

Building a comprehensive observability stack requires implementing monitoring, logging, and tracing systems using tools like Prometheus, Grafana, ELK Stack, and Jaeger to proactively identify and resolve performance bottlenecks.

What is the best way to implement distributed tracing for enterprise-scale applications?

Distributed tracing for enterprise-scale applications is best implemented using Jaeger, Zipkin, and OpenTelemetry, providing deep application insight and enabling effective debugging across complex distributed system architectures.

How do I set up SLI and SLO management for reliability engineering?

SLI and SLO management for reliability engineering involves defining service level indicators and objectives within your observability strategy, integrating incident response workflows and automated runbooks to maintain application performance.

Can I use Prometheus and Grafana for centralized log management and analysis?

Prometheus and Grafana primarily handle monitoring and metrics visualization. For centralized log management and analysis, you should deploy the ELK Stack, Loki, or Splunk to aggregate and analyze application logs.

Does this observability approach work with cloud-native monitoring tools and DataDog?

This observability approach works with cloud-native monitoring tools and DataDog, allowing you to implement comprehensive monitoring solutions tailored to your specific enterprise-scale application reliability requirements.

How do I design an incident response workflow with automated alerting?

Designing an incident response workflow with automated alerting involves configuring PagerDuty integration and creating automated runbooks within your observability strategy to handle performance bottlenecks and system failures.