observability-engineer

Design monitoring, logging, and tracing systems with SLI/SLO management.

1|Updated Feb 6, 2026
One-click install
npx skills add https://github.com/Adam-Guerin/Asmblr --skill observability-engineer-adam-guerin
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability-engineer
Source: https://github.com/Adam-Guerin/Asmblr/tree/main/skills/observability-engineer
Command: npx skills add https://github.com/Adam-Guerin/Asmblr --skill observability-engineer-adam-guerin

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill addresses the critical need for robust monitoring, logging, and tracing systems to ensure the reliability, performance, and stability of enterprise-scale applications.

Core Features & Use Cases

  • Comprehensive Observability: Design and implement end-to-end solutions for monitoring, logging, and tracing.
  • SLI/SLO Management: Define, track, and manage Service Level Indicators and Objectives to ensure service quality.
  • Incident Response: Establish workflows for efficient and effective incident detection, diagnosis, and resolution.
  • Use Case: You need to set up a complete observability stack for a new microservices-based e-commerce platform, including defining SLIs for critical user journeys and configuring alerts for potential performance degradations.

Quick Start

Design a comprehensive monitoring strategy for a microservices architecture with 50+ services.

Frequently Asked Questions about observability-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I build a production-ready observability stack for microservices?

Build a production-ready observability stack by integrating monitoring, logging, and tracing systems using tools like Prometheus, Grafana, Jaeger, ELK Stack, and OpenTelemetry to ensure enterprise-scale application reliability.

What's the best way to define and manage SLI and SLO for service reliability?

Define and manage SLI and SLO by establishing Service Level Indicators and Objectives to track service quality, implementing SRE practices to ensure reliability, and configuring alerts for potential performance degradations.

How do I set up incident response workflows for diagnosing application performance issues?

Set up incident response workflows by establishing processes for efficient incident detection, diagnosis, and resolution, leveraging comprehensive tracing and logging data to diagnose performance issues in enterprise-scale applications.

Can I use OpenTelemetry and Jaeger together for distributed tracing in a large architecture?

Use OpenTelemetry and Jaeger together to implement distributed tracing across large architectures, enabling end-to-end visibility into service interactions and diagnosing latency issues within comprehensive observability strategies.

What is needed to implement a comprehensive monitoring strategy for 50+ services?

Implementing a monitoring strategy for 50+ services requires designing comprehensive observability systems that integrate Prometheus and Grafana, managing SLI/SLO metrics, and establishing incident response workflows for enterprise-scale applications.