What problem does it solve?
This Skill addresses the critical need for defining, monitoring, and managing the reliability and performance of Urbit deployments, ensuring they meet agreed-upon service levels.
Core Features & Use Cases
- Define SLAs: Establish clear Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for availability, latency, and error rates.
- Monitor Compliance: Implement checks and balances to track performance against defined SLOs.
- Incident Management: Outline a structured process for responding to and resolving incidents based on severity.
- Error Budgeting: Track downtime and performance degradation against a defined error budget.
- Use Case: A system administrator can use this skill to set up a 99.9% uptime target for their Urbit ship, define how to measure it, and establish penalties for breaches, ensuring consistent service delivery.
Quick Start
Use the sla-management skill to define a 99.9% monthly availability SLO for your Urbit ship.