observability-engineer

Design and implement monitoring, logging, and distributed tracing systems.

Updated Jul 6, 2026
One-click install
npx skills add https://github.com/brenordv/claude-skill-set --skill observability-engineer-brenordv
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability-engineer
Source: https://github.com/brenordv/claude-skill-set/tree/main/skills/observability-engineer
Command: npx skills add https://github.com/brenordv/claude-skill-set --skill observability-engineer-brenordv

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This skill addresses the complexity of maintaining system reliability by providing a structured framework for monitoring, logging, and tracing, preventing production outages and performance degradation.

Core Features & Use Cases

  • SLI/SLO Management: Define and track service level indicators and objectives to ensure business-aligned reliability.
  • Observability Architecture: Design comprehensive monitoring stacks including distributed tracing, log aggregation, and alerting.
  • Use Case: When a microservice experiences intermittent latency, use this skill to correlate distributed traces with infrastructure metrics to identify the root cause and configure an automated alert for future occurrences.

Quick Start

Use the observability-engineer skill to design a monitoring strategy and define SLIs for the new user authentication service.

Frequently Asked Questions about observability-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design a distributed tracing architecture for microservices?

Design distributed tracing architectures by implementing OpenTelemetry standards to correlate infrastructure metrics with logs. This framework provides a structured monitoring stack to prevent production outages and identify intermittent latency root causes across enterprise-scale applications.

How do I define SLI and SLO management for a new service?

Define SLI and SLO management by tracking service level indicators aligned with business reliability objectives. This structured framework ensures you measure and maintain production-grade reliability standards for new user authentication services and other critical enterprise workloads.

What is the best way to configure actionable alerting strategies?

Configure actionable alerting strategies by correlating distributed traces with infrastructure metrics to trigger automated alerts. This approach ensures future performance degradations are caught early while satisfying requirements for cost-effective data retention.

Can I use OpenTelemetry standards for enterprise-scale observability?

Yes, OpenTelemetry standards are fully supported for designing enterprise-scale observability systems. The framework satisfies requirements for OpenTelemetry compliance while implementing monitoring, logging, and distributed tracing across complex microservice architectures.

Why do I need observability for incident response workflows?

Observability is needed for incident response workflows to correlate distributed traces with infrastructure metrics and identify root causes. This structured framework prevents production outages by providing actionable alerting and performance regression analysis during incidents.

Does this approach handle cost-effective data retention for log aggregation?

Yes, the observability architecture handles cost-effective data retention while implementing comprehensive log aggregation and monitoring. It ensures enterprise-scale applications maintain production-grade reliability without excessive storage costs for metrics and trace data.