incident-response-ops

Coordinate structured incident response for AI SaaS platforms with severity levels and postmortem templates.

1|Updated Mar 12, 2026
One-click install
npx skills add https://github.com/XiaoPuOuO/VFactory --skill incident-response-ops
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-response-ops
Source: https://github.com/XiaoPuOuO/VFactory/tree/main/paperclip-official/AgentSetting/skills/incident-response-ops
Command: npx skills add https://github.com/XiaoPuOuO/VFactory --skill incident-response-ops

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill provides a structured, blameless incident response framework for AI SaaS / control-plane systems, enabling rapid triage, clear roles, and auditable timelines to reduce downtime and coordination overhead.

Core Features & Use Cases

  • Structured incident workflow with severity levels, canonical timeline, and escalation guards.
  • Role-based coordination (Incident Commander, Operations Lead, Tech Lead, Communications Lead, Scribe) with explicit ownership and cadence.
  • Postmortem / RCA templates, acceptance criteria for closure, and rollback/mitigation guidance for fast recovery.
  • Evidence preservation and timeline logging to support audits and learning.

Quick Start

Create an incident, assign an Incident Commander, and start the canonical timeline using the default severity model.

Frequently Asked Questions about incident-response-ops

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I coordinate blameless incident response for an AI SaaS platform outage?

Blameless incident response is coordinated through a structured workflow with severity levels, role-based assignments, and a canonical timeline. This framework reduces downtime by enforcing clear ownership and escalation guards for AI control-plane systems.

What roles are needed for large-scale SaaS incident response?

Large-scale SaaS incident response requires an Incident Commander, Operations Lead, Tech Lead, Communications Lead, and Scribe. These roles provide explicit ownership and cadence to manage degraded services and coordinate rapid triage.

How do I create a postmortem or RCA after a service outage?

To create a postmortem or RCA after a service outage, use structured templates with acceptance criteria for closure. This ensures blameless documentation of root causes, rollback mitigations, and auditable timelines for future learning.

When do I need to use rollback criteria during incident management?

Rollback criteria during incident management are needed when a deployment causes degraded services or outages. The structured workflow enforces these criteria to guide fast recovery and mitigate service interruptions in AI control-plane systems.

Can this incident response workflow be used for security events on DevOps on-call rotations?

Yes, this incident response workflow applies to DevOps on-call rotations handling security events. It provides a fixed workflow with timeline logging and evidence preservation to support audits across large-scale SaaS platforms.

How does timeline logging support blameless postmortems?

Timeline logging supports blameless postmortems by preserving evidence and creating an auditable canonical timeline of incident events. This structured record enables accurate root cause analysis and learning without assigning blame.