oncall-sre-operations

Automate on-call incident triage, severity routing, runbook selection, and escalation for AI SaaS services.

1|Updated Mar 12, 2026
One-click install
npx skills add https://github.com/XiaoPuOuO/VFactory --skill oncall-sre-operations
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: oncall-sre-operations
Source: https://github.com/XiaoPuOuO/VFactory/tree/main/paperclip-official/AgentSetting/skills/oncall-sre-operations
Command: npx skills add https://github.com/XiaoPuOuO/VFactory --skill oncall-sre-operations

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill standardizes on-call incident response for AI SaaS by unifying alert intake, triage, escalation, and handoff processes under time pressure.

Core Features & Use Cases

  • Structured alert intake: Normalize incoming signals into a single incident candidate with required fields (source, service, environment, trigger, first seen, status, user impact, and known changes).
  • Severity routing & runbook mapping: Assign a precise severity and select the closest runbook class for remediation.
  • Escalation & ownership management: Enforce escalation policies and assign incident commanders, owners, and communication roles.
  • Handoff documentation & postmortem readiness: Produce a handoff-safe record with current status, mitigations, owners, and next steps.

Quick Start

Initiate the on-call incident triage flow upon alert receipt and populate severity, runbook, and ownership fields to trigger escalation and handoff.

Frequently Asked Questions about oncall-sre-operations

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate on-call incident triage for production outages?

Automate on-call incident triage by normalizing incoming alerts into a single incident candidate with required fields like source, service, environment, and user impact. This standardizes alert intake under time pressure during production outages.

What is the best way to map runbooks to incident severity during an outage?

Runbook mapping assigns a precise severity to the incident and selects the closest runbook class for remediation. This ensures correct severity routing and runbook selection during degraded reliability or outages.

How does incident escalation and ownership management work for SRE teams?

Incident escalation enforces predefined policies to assign incident commanders, owners, and communication roles. This manages ownership and ensures proper escalation across production systems during service disruptions.

How do I document structured handoff data for on-call transitions?

Structured handoff documentation produces a handoff-safe record containing current status, mitigations, owners, and next steps. This ensures seamless on-call transitions and maintains postmortem readiness for AI SaaS services.

Can I use this incident response process for AI SaaS services specifically?

Yes, this incident response process is designed specifically for AI SaaS services. It applies to alert intake, severity routing, runbook selection, escalation, and handoff scenarios across production systems during outages.

What is needed to start the on-call incident triage flow?

To start the on-call incident triage flow, you need an incoming alert signal. Upon alert receipt, populate severity, runbook, and ownership fields to trigger escalation and structured handoff.