investigate-alert

Diagnose production alerts by aggregating logs, metrics, traces, and source code.

11|Updated Sep 6, 2011
One-click install
npx skills add https://github.com/tclem/dotfiles --skill investigate-alert-tclem
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: investigate-alert
Source: https://github.com/tclem/dotfiles/tree/main/copilot/skills/investigate-alert
Command: npx skills add https://github.com/tclem/dotfiles --skill investigate-alert-tclem

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

解决在遇到监控、日志、追踪或指标检测到的警报时,快速定位根因和分析证据的难题。

Core Features & Use Cases

  • 故障诊断:分析各种监控和问题报告,识别警报的潜在原因。
  • 证据收集:整合指标、日志、追踪和代码信息,为排查提供详实依据。
  • 使用场景:当系统出现性能下降或故障时,快速聚合相关数据并生成调查报告。

Quick Start

输入警报信息或监控链接,启动自动诊断流程,生成详细调查报告。

Frequently Asked Questions about investigate-alert

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I investigate production alerts and identify root causes?

Production alert investigation aggregates logs, metrics, traces, and source code context to identify root causes. By analyzing telemetry data and dependency checks, it formulates hypotheses for effective incident resolution.

What is the best way to troubleshoot system anomalies detected via monitoring tools?

The best way to troubleshoot anomalies detected via monitoring tools is by analyzing telemetry data and source code context. Aggregating logs and traces ensures comprehensive evidence collection for diagnosing the issue.

How does trace analysis help with incident investigation in production environments?

Trace analysis assists incident investigation by providing request-level visibility across dependencies. Integrating traces with metrics and logs enables comprehensive evidence collection to pinpoint the exact root cause of production alerts.

Can I diagnose performance degradation using only log analysis and metrics?

While log analysis and metrics are useful, diagnosing performance degradation is more effective when combined with trace analysis and source code context. Aggregating all telemetry data ensures comprehensive evidence collection for accurate troubleshooting.

When do I need to aggregate logs and code context for alert troubleshooting?

You need to aggregate logs and code context for alert troubleshooting when monitoring tools detect anomalies or incidents occur. This comprehensive evidence collection is essential for formulating accurate incident resolution hypotheses.

Does incident investigation work without integrating source code context?

Incident investigation without source code context is incomplete for root cause analysis. Integrating code context with telemetry data and traces ensures comprehensive evidence collection and accurate hypothesis formulation for resolution.