What problem does it solve?
Teams often enter production with vague expectations like "it should be fast and reliable" but no measurable targets. This Skill turns those expectations into concrete SLIs, SLOs, error budgets, load models, and a reliability backlog that can be reviewed and executed.
Core Features & Use Cases
- SLO and Error Budget Definition: Derives measurable SLIs (latency p50/p95/p99, availability, error rate, false-positive rate) with targets, windows, and error budgets anchored to PRD NFR ids.
- Load Modeling and Test Matrix: Defines steady, peak, spike, and soak profiles and maps them to critical user journeys with pass/fail thresholds for Azure Load Testing execution.
- Reliability Backlog and Observability Hooks: Produces a prioritized SRE backlog and names the metrics and traces needed to measure each SLI in production.
- Use Case: Before launching an incident-response app, use this Skill to convert PRD NFRs like "Critical ≤ 60s" into a full performance plan with load profiles, capacity assumptions, and degradation behavior.
Quick Start
Ask the agent to create a performance and SLO plan for your application using its critical user journeys and traffic assumptions.