devops-expert

Review CI/CD pipelines and platform designs for reliability and operational excellence.

1|Updated Apr 13, 2026
One-click install
npx skills add https://github.com/felixgeelhaar/skills --skill devops-expert-felixgeelhaar
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: devops-expert
Source: https://github.com/felixgeelhaar/skills/tree/main/devops-expert
Command: npx skills add https://github.com/felixgeelhaar/skills --skill devops-expert-felixgeelhaar

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps teams design and operate reliable, scalable production systems by applying advanced DevOps thinking, including SRE principles, GitOps, IaC, and platform engineering practices, across CI/CD pipelines, incident management, observability, and governance.

Core Features & Use Cases

  • SRE-leaning reviews of pipelines, SLOs/SLIs, error budgets, and incident response for faster recovery and reduced toil.
  • Platform design guidance to reduce cognitive load on stream-aligned teams through internal developer platforms, shared telemetry, and standard patterns.
  • Operational pairing across engineering, security, and product disciplines to validate trade-offs and align architectural decisions with business goals.

Quick Start

Describe your CI/CD architecture and I will evaluate it for reliability and operational excellence.

Frequently Asked Questions about devops-expert

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLOs and error budgets for my CI/CD pipelines?

To define SLOs and error budgets for CI/CD pipelines, you establish SLIs measuring reliability gaps, then apply SRE principles to align deployment strategies with operational excellence and faster recovery targets.

What's the best way to reduce cognitive load for platform engineering teams?

The best way to reduce cognitive load in platform engineering is designing internal developer platforms with shared telemetry, standard patterns, and clear platform boundaries that validate architectural trade-offs across teams.

How does GitOps improve infrastructure reviews and IaC reliability?

GitOps improves infrastructure reviews and IaC reliability by applying operational excellence practices to evaluate deployment strategies, validate trade-offs, and resolve reliability gaps across production systems adopting GitOps workflows.

When do I need a postmortem framework for incident management?

You need a postmortem framework for incident management when adopting SRE principles to reduce toil, requiring structured reviews of error budgets, SLIs, and incident response to achieve faster recovery and operational excellence.

Can I use this approach for cross-functional operational pairing across security and product teams?

Yes, you can use this approach for operational pairing across engineering, security, and product disciplines to validate trade-offs, satisfy governance requirements, and align architectural decisions with business goals.

Why does my observability strategy fail to resolve reliability gaps across teams?

Your observability strategy fails to resolve reliability gaps when it lacks shared telemetry, standard patterns, and clear platform boundaries needed to reduce cognitive load and align operational excellence across stream-aligned teams.