observability-stack

Standardize metrics, logs, and traces across Prometheus, Loki, Tempo, and Grafana.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/matt-metivier/zk-hub --skill observability-stack-matt-metivier
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability-stack
Source: https://github.com/matt-metivier/zk-hub/tree/main/skills/general/tools/observability-stack
Command: npx skills add https://github.com/matt-metivier/zk-hub --skill observability-stack-matt-metivier

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Observability is often fragmented across tools and teams; this skill consolidates best practices for instrumenting, collecting, and presenting metrics, logs, and traces to reduce toil and speed incident response.

Core Features & Use Cases

  • Standardized instrumentation guidelines for metrics, logs, traces, and profiling.
  • Ready-made Grafana dashboards and SLO alert patterns for unified visibility.
  • Use Case: a services team can observe cross-service latency and error budgets with a single source of truth.

Quick Start

Configure your environment to adopt the observability stack and instrument a sample app using the provided dashboards and rules.

Frequently Asked Questions about observability-stack

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I unify observability patterns across Prometheus, Loki, and Tempo?

Standardizing observability patterns across Prometheus and Grafana provides ready-made dashboards and SLO alert patterns. This enforces consistent instrumentation conventions for metrics, logs, and traces to reduce toil and misconfigurations.

What is the best way to standardize metrics and logs for multi-service architectures?

Standardizing metrics and logs for multi-service architectures requires consistent instrumentation conventions across services. This provides unified visibility into cross-service latency and error budgets to speed up incident response.

How do I set up SLO alert patterns and Grafana dashboards for reliable monitoring?

Setting up SLO alert patterns and Grafana dashboards involves applying ready-made templates for unified visibility. These pre-configured resources enforce instrumentation conventions and track error budgets across multi-service environments.

Can I use this observability stack for cross-service telemetry in a SaaS environment?

Yes, this observability stack is designed for multi-service architectures and SaaS environments. It includes integration guidance for cross-tool telemetry, ensuring consistent metrics, logs, and traces are collected for reliable monitoring.

Why does fragmented observability across tools increase incident response toil?

Fragmented observability across tools increases incident response toil because inconsistent metrics, logs, and traces prevent a single source of truth. Standardizing instrumentation and cross-tool telemetry collection reduces misconfigurations and speeds up resolution.