observability

Correlate logs, metrics, and traces across multiple backends to identify root causes.

654|77|Updated Jan 20, 2026
One-click install
npx skills add https://github.com/incidentfox/incidentfox --skill observability-incidentfox
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability
Source: https://github.com/incidentfox/incidentfox/tree/main/sre-agent/.claude/skills/observability
Command: npx skills add https://github.com/incidentfox/incidentfox --skill observability-incidentfox

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill provides a structured approach to observability analysis across logs, metrics, and traces, helping teams identify root causes and prioritize fixes by correlating signals from multiple backends.

Core Features & Use Cases

  • Cross-backend signal correlation: analyze logs, metrics, and traces across Coralogix, Datadog, Honeycomb, Splunk, Elasticsearch/OpenSearch, and Jaeger.
  • Anomaly detection guidance: identify unusual patterns, spikes, and failure clusters in production systems.
  • Investigation playbook: step-by-step methodology to triage incidents using aggregated statistics and targeted sampling.
  • Use Case: When investigating a high latency incident, align traces with metrics to locate bottlenecks and dependency issues.

Quick Start

Use the observability skill to begin a guided analysis of the attached log stream and metrics dashboard.

Frequently Asked Questions about observability

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I correlate logs, metrics, and traces to find a root cause during incident response?

Correlating logs, metrics, and traces requires a structured analysis workflow that aligns signals across backends to identify root causes. This skill guides teams through incident response by applying recommended queries and sampling strategies to pinpoint bottlenecks.

Does this observability analysis approach work with Datadog, Splunk, and Elasticsearch?

Yes, this observability analysis approach works with Datadog, Splunk, Elasticsearch, and OpenSearch. It also supports Coralogix, Honeycomb, and Jaeger, allowing teams to perform cross-backend signal correlation during on-call investigations.

What is the best way to investigate high latency incidents using traces and metrics?

The best way to investigate high latency incidents is to align traces with metrics to locate bottlenecks and dependency issues. This skill provides a step-by-step investigation playbook using aggregated statistics and targeted sampling to triage production anomalies.

How do I detect anomalies and failure clusters in production systems?

Detecting anomalies and failure clusters in production systems involves analyzing observability data for unusual patterns and spikes. This skill provides anomaly detection guidance by correlating logs, metrics, and traces to identify failure clusters and prioritize fixes.

Can I use this methodology for post-mortem analysis after an on-call incident?

Yes, you can use this methodology for post-mortem analysis after an on-call incident. It applies a structured observability workflow to correlate signals from multiple backends, helping teams review aggregated statistics and document root causes effectively.

How do I triage incidents using aggregated statistics and targeted sampling?

Triage incidents using aggregated statistics and targeted sampling by following a structured investigation playbook. This skill enforces a step-by-step methodology to query observability data across logs, metrics, and traces, ensuring effective root cause identification.