observabilityaudit

Audit end-to-end telemetry coverage across logs, traces, metrics, dashboards, and runbooks.

10|3|Updated May 26, 2026
One-click install
npx skills add https://github.com/agentik-os/OmegaOS --skill observabilityaudit
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observabilityaudit
Source: https://github.com/agentik-os/OmegaOS/tree/main/skills/audits/observabilityaudit
Command: npx skills add https://github.com/agentik-os/OmegaOS --skill observabilityaudit

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Forensic observability audit to verify end-to-end telemetry coverage across logs, traces, metrics, dashboards, and runbooks, ensuring you can diagnose incidents in production with confidence.

Core Features & Use Cases

  • 18-phase observability audit framework covering logging, tracing, metrics, dashboards, alerts, and SLOs
  • Identify instrumentation gaps, data leakage, and end-to-end traceability issues
  • Produce actionable verdicts, fix plans, and progress tracking across teams

Quick Start

Run the observability audit against the project to generate a complete telemetry coverage report.

Frequently Asked Questions about observabilityaudit

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I audit production telemetry coverage for critical service paths?

A comprehensive observability audit evaluates end-to-end telemetry coverage across logs, traces, metrics, dashboards, and runbooks to verify you can diagnose production incidents with confidence. It produces actionable verdicts and a fix plan.

What is end-to-end observability traceability for authentication and payment mutations?

End-to-end observability traceability verifies that complex service mutations like authentication and payments have complete distributed tracing, structured logging, and metrics collection. This ensures you can diagnose production incidents across critical paths.

How do I identify instrumentation gaps and telemetry data leakage in my systems?

You identify instrumentation gaps and data leakage by conducting an observability audit that checks structured logging, distributed tracing, metrics collection, alerting rules, and SLO/SLI definitions. It outputs a formal evidence package with actionable verdicts.

Can I generate runbook-ready fix plans and track progress for SLO and SLI definitions?

Yes, an observability audit produces runbook-ready verdicts, a fix plan, and progress tracking across teams. It verifies your SLO and SLI definitions alongside alerting rules to ensure complete production telemetry coverage.

Does the observability audit work for complex services with critical paths requiring 18 phases of verification?

Yes, the observability audit is designed for complex services with critical paths requiring end-to-end traceability. It conducts verification across 18 phases, covering logs, traces, metrics, dashboards, alerts, and SLOs.

What are the limitations of relying only on logging and metrics without distributed tracing?

Relying only on logging and metrics creates end-to-end traceability issues for complex service paths. A complete observability audit requires verifying distributed tracing, SLO definitions, and alerting rules to prevent untraceable production incidents.