ops-troubleshooting

Analyze Nightingale alerts, metrics, logs, and host details to diagnose incidents.

13.2k|1.8k|Updated Mar 3, 2020
One-click install
npx skills add https://github.com/ccfos/nightingale --skill ops-troubleshooting
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ops-troubleshooting
Source: https://github.com/ccfos/nightingale/tree/main/aiagent/skill/embedded/builtin/ops-troubleshooting
Command: npx skills add https://github.com/ccfos/nightingale --skill ops-troubleshooting

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill streamlines the process of fault localization and root cause analysis within the Nightingale monitoring platform, reducing downtime and improving system stability.

Core Features & Use Cases

  • Incident Analysis: Analyze specific alert details, metrics trends, logs, and host states to identify causes of failures.
  • Holistic Troubleshooting: Conduct targeted investigations across alerts, data sources, and logs with minimal manual effort.
  • Use Case: A system administrator detects a spike in alerts and uses this Skill to trace back to the underlying metrics and logs, pinpointing the faulty service or host quickly.

Quick Start

Provide the current alert ID or relevant target information to begin diagnostics and gather detailed insights on the incident.

Frequently Asked Questions about ops-troubleshooting

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I perform root cause analysis for a Nightingale monitoring alert?

Perform root cause analysis by providing the alert ID to query metrics trends, logs, and host states. This identifies the underlying faulty service or host quickly with minimal manual effort.

What is the best way to troubleshoot a system incident in Nightingale?

Troubleshoot a system incident by running targeted investigations across alerts, data sources, and logs. The tool traces back from alert spikes to underlying metrics and host details to pinpoint failures.

How does incident response work with Nightingale metrics and logs?

Incident response works by automating the query and analysis of metrics, logs, and host details. It correlates alert details with system data to support incident investigation workflows and ensure faster system recovery.

Do I need an alert ID to start system troubleshooting in Nightingale?

You need the current alert ID or relevant target information to begin diagnostics. Providing this target information gathers detailed insights on the incident and initiates the automated fault localization process.

Can I analyze host details and data sources when diagnosing Nightingale platform issues?

You can analyze host details and data sources when diagnosing Nightingale platform issues. The analysis leverages internal data queries to conduct holistic troubleshooting across alerts, metrics, and logs.