system-monitoring-expert

Design monitoring dashboards, telemetry schemas, and alerting strategies for platform health.

1|1|Updated Mar 22, 2026
One-click install
npx skills add https://github.com/zzafergok/skills --skill system-monitoring-expert
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: system-monitoring-expert
Source: https://github.com/zzafergok/skills/tree/main/08-engineering-quality/system-monitoring-expert
Command: npx skills add https://github.com/zzafergok/skills --skill system-monitoring-expert

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Modern software platforms struggle with visibility into health, performance, and reliability. This skill provides a structured approach to design and implement monitoring dashboards, telemetry schemas, and alerting that translate complex operations into actionable insights.

Core Features & Use Cases

  • KPI hierarchy planning for platform health, reliability, and SLOs.
  • Telemetry schema design, event taxonomy, and admin dashboard patterns.
  • Alert design, incident response readiness, and observability best practices.
  • Use Case: When building or evolving monitoring for APIs, services, or platforms, apply this skill to align dashboards with business and technical KPIs.

Quick Start

Configure an initial monitoring dashboard by selecting platform health KPIs, define telemetry events, and outline an alerting scheme for critical routes.

Frequently Asked Questions about system-monitoring-expert

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design a monitoring dashboard for microservices and cloud components?

Design monitoring dashboards by establishing a KPI hierarchy for platform health, defining telemetry schemas, and outlining alerting schemes for critical API routes. This structured approach aligns health visualizations with business and technical SLOs across microservices and cloud environments.

What is the best way to structure telemetry events and schemas for observability?

Structure telemetry events by defining a clear event taxonomy and schema that supports data quality and security considerations. Applying observability best practices ensures your telemetry translates complex platform operations into actionable insights for admin dashboards.

How do I create an alerting strategy for API routes and production environments?

Create an alerting strategy by defining incident response readiness and mapping alerts to critical API routes and platform KPIs. This approach ensures production environments maintain reliability and provides structured admin UI patterns for managing platform health.

Can I use this approach to evolve existing platform health visualizations and KPIs?

Yes, you can evolve existing platform health visualizations by redefining KPI hierarchies for reliability and SLOs. This structured approach helps align current dashboards with business and technical KPIs, ensuring effective visibility into system health as your platform scales.

When do I need a structured approach to observability and admin dashboard patterns?

You need a structured approach to observability when modern software platforms struggle with visibility into health, performance, and reliability. Defining requirements for telemetry schemas, alerting strategies, and admin UI patterns translates complex operations into actionable insights.