talos-cluster-ops

Correlate Talos, Kubernetes, Cilium, Longhorn, and GitOps evidence to diagnose cluster failures.

3|1|Updated Dec 3, 2025
One-click install
npx skills add https://github.com/Probably-Group/Dev-AID --skill talos-cluster-ops
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: talos-cluster-ops
Source: https://github.com/Probably-Group/Dev-AID/tree/main/.dev-aid/skills/expert/talos-cluster-ops
Command: npx skills add https://github.com/Probably-Group/Dev-AID --skill talos-cluster-ops

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

It resolves slow, broken, or unstable Talos-based Kubernetes clusters by guiding investigation across hardware, Talos, Kubernetes, Cilium networking, Longhorn storage, GitOps drift, and security signals, then proposing remediation that requires explicit approval.

Core Features & Use Cases

  • Autonomous investigation workflows for pod crashes, node failures, networking symptoms, storage problems, and GitOps sync failures with structured, multi-phase evidence gathering.
  • Approval-gated remediation with explicit risk levels, rollback expectations, and optional snapshot requirements to reduce operational mistakes.
  • Cross-layer correlation using Talos health/logs, Kubernetes events/logs/metrics, Cilium/Hubble flow evidence, Longhorn replica health, ArgoCD drift checks, and security signals (Falco/Tetragon/Trivy/SPIRE).

Quick Start

Tell your AI to investigate a failing workload with: "Investigate why my pod rabbitmq-0 in production is crashing using talos-cluster-ops and propose the safest remediation I can approve."

Frequently Asked Questions about talos-cluster-ops

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I troubleshoot CrashLoopBackOff and ImagePullBackOff errors in a Talos Linux Kubernetes cluster?

You diagnose Pending or Error pod states in Talos Kubernetes environments through a structured multi-phase investigation that gathers evidence across Kubernetes events, Longhorn storage health, and Cilium networking to identify scheduling or resource constraints.

What is the best way to diagnose Node NotReady and node pressure conditions on Talos Linux?

Diagnosing Node NotReady and pressure conditions on Talos Linux requires cross-layer correlation of Talos health metrics, etcd control-plane status, and security telemetry to pinpoint hardware or system-level stressors before applying approved fixes.

How do I resolve ArgoCD sync mismatches and GitOps drift in a Talos Kubernetes environment?

Resolving ArgoCD sync mismatches and GitOps drift in a Talos Kubernetes environment involves checking ArgoCD state against cluster realities, cross-referencing Talos and Kubernetes telemetry, and proposing risk-assessed remediation actions requiring explicit approval.

Does this Kubernetes troubleshooting approach work with Cilium, Hubble, and Longhorn storage?

Yes, this troubleshooting approach works directly with Cilium and Hubble networking evidence, Longhorn replica health checks, and ArgoCD drift detection to comprehensively diagnose and resolve complex Talos-based cluster failures.

Why does my Talos Kubernetes rollout get stuck and how can I safely remediate it?

Stuck rollouts in Talos Kubernetes clusters are remediated safely by autonomously investigating cross-layer evidence from Talos to Longhorn, then requiring explicit user approval with defined risk levels and rollback expectations before applying changes.