What problem does it solve?
This Skill provides a shared validation and monitoring checklist for service deployment and upgrades, helping teams ensure a consistent, reliable rollout process.
Core Features & Use Cases
- Deployment Validation: after deploying or upgrading a service, verify pods are running, logs are clean, Argo CD shows Synced/Healthy, ingress is accessible, and dependent services remain operational.
- Monitoring Requirements: ensure metrics endpoints exist, configure Prometheus monitoring, and verify metrics appear.
- Alerting & Dashboards: confirm alert rules, dashboards, and accessibility.
- Human-in-the-Loop Workflow: AI Assistant MUST perform changes, show diffs, stage changes, and guide the user through approvals.
- Anti-Patterns: avoid risky practices like pushing without approval or skipping monitoring.
Quick Start
Use this checklist to validate a newly deployed service by stepping through each item and recording the results.