observability-architecture

Design observability architecture for metrics, logs, traces, dashboards, and alert routing.

Updated Mar 18, 2026
One-click install
npx skills add https://github.com/Canepro/codex-skills --skill observability-architecture-canepro
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability-architecture
Source: https://github.com/Canepro/codex-skills/tree/main/skills/observability-architecture
Command: npx skills add https://github.com/Canepro/codex-skills --skill observability-architecture-canepro

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Design or review observability architecture across metrics, logs, traces, dashboards, alert routing, telemetry standards, retention, and ownership. Use when the user wants a durable monitoring strategy, telemetry platform design, signal governance, cost or retention trade-offs, or an observability stack review rather than live alert triage.

Core Features & Use Cases

  • Recommend durable telemetry-system design and governance.
  • Standardize metrics/logs/traces, alert routing, and dashboard ownership.
  • Evaluate cost, retention, and signal quality trade-offs across teams.

Quick Start

Follow the steps to define the observation architecture while aligning on ownership and cost constraints.

Frequently Asked Questions about observability-architecture

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design an observability architecture for metrics, logs, and traces?

Observability architecture design establishes durable monitoring across metrics, logs, traces, dashboards, and alert routing. It defines a signal model, collection engineering guidelines, retention policies, and clear ownership standards across teams.

What is the best way to standardize telemetry and signal governance across teams?

Telemetry governance standardizes metrics, logs, and traces through a defined signal model and collection guidelines. Cross-team reviews establish ownership, dashboard standards, and alert routing decisions to maintain signal quality and durability.

How do I plan an observability platform migration without losing signal coverage?

Observability platform migrations require defining a signal model and collection engineering guidelines before moving dashboards and alert routing. Scoping cross-team reviews ensures retention policies and ownership are established to maintain continuous monitoring coverage.

How do I evaluate cost and retention trade-offs for telemetry data?

Evaluating cost and retention trade-offs involves reviewing telemetry standards and signal quality across teams. The architecture defines retention policies and signal governance to balance monitoring durability against storage and operational costs.

When do I need to define alert routing decisions and dashboard ownership?

Alert routing decisions and dashboard ownership are needed when establishing observability architecture or during platform migrations. Defining these standards ensures durable monitoring, clear accountability, and consistent signal governance across engineering teams.

Does observability architecture design work for live alert triage and incident response?

Observability architecture design does not handle live alert triage or incident response. It focuses on durable monitoring strategy, telemetry platform design, signal governance, and retention trade-offs rather than real-time alert resolution.