observability-engineer

Develop observability systems with SLI/SLO management and incident response workflows.

Updated Dec 18, 2025
One-click install
npx skills add https://github.com/JesusFigueroa25/SEABOT --skill observability-engineer-jesusfigueroa25
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability-engineer
Source: https://github.com/JesusFigueroa25/SEABOT/tree/main/PROYECTO/backend-seabot/.codex/skills/observability-engineer
Command: npx skills add https://github.com/JesusFigueroa25/SEABOT --skill observability-engineer-jesusfigueroa25

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill empowers you to build and manage complex observability systems for production environments, with a focus on SLI/SLO management and incident response.

Core Features & Use Cases

  • Monitoring & Metrics Infrastructure: Implement Prometheus, Grafana, InfluxDB, and more.
  • Distributed Tracing & APM: Deploy Jaeger, Zipkin, and AWS X-Ray.
  • Log Management & Analysis: Work with ELK Stack, Fluentd, and Splunk.

Quick Start

Execute 'setup-observability' to create a monitoring strategy with dashboards and alerts aligned to SLOs.

Frequently Asked Questions about observability-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I set up SLI and SLO management for enterprise production monitoring?

SLI and SLO management for production monitoring is established by executing 'setup-observability' to generate a monitoring strategy with dashboards and alerts aligned to SLOs.

What's the best way to implement distributed tracing and log management together?

Implement distributed tracing and log management together by deploying APM tools like Jaeger or Zipkin alongside log analysis stacks such as the ELK Stack or Splunk.

Can I use Prometheus and Grafana to build incident response workflows?

Prometheus and Grafana can be used to build incident response workflows by integrating them into a comprehensive observability system that triggers alerts based on SLO thresholds.

How does observability with AWS X-Ray work for enterprise-scale applications?

Observability with AWS X-Ray for enterprise-scale applications works by providing distributed tracing capabilities that map requests across services to identify latency and bottlenecks.

Do I need SRE practices to deploy Fluentd and InfluxDB for production monitoring?

SRE practices are required to effectively deploy Fluentd and InfluxDB, as this Skill requires expertise in SRE practices and modern observability stacks for production environments.

Why does my observability strategy need SLO-aligned alerts instead of standard thresholds?

An observability strategy needs SLO-aligned alerts instead of standard thresholds to ensure monitoring focuses on user-facing reliability and incident response workflows rather than isolated system metrics.