llm-failure-modes

Identify and mitigate LLM cognitive failure modes in high-stakes reasoning.

1|Updated Feb 24, 2026
One-click install
npx skills add https://github.com/dzackgarza/ai --skill llm-failure-modes
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: llm-failure-modes
Source: https://github.com/dzackgarza/ai/tree/main/opencode/skills/llm-failure-modes
Command: npx skills add https://github.com/dzackgarza/ai --skill llm-failure-modes

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Helps teams identify, analyze, and mitigate cognitive failure modes in LLM-driven reasoning, reducing errors in high-stakes decisions.

Core Features & Use Cases

  • Failure mode taxonomy: Provides a structured taxonomy of common cognitive failures in LLMs.
  • Editorial guidelines: Offers checklists and best practices for evaluating model outputs.
  • Practical validation: Enables post-hoc analysis and design of guardrails for reliable reasoning in engineering contexts.
  • Use Case: In critical data science tasks, use this guide to audit the model's reasoning path and surface potential missteps.

Quick Start

Apply the failure-mode checklist to the model's response and record any flagged issues for corrective action.

Frequently Asked Questions about llm-failure-modes

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
What are common LLM cognitive failure modes in high-stakes reasoning?

LLM cognitive failure modes encompass reasoning errors and cognitive blind spots that degrade high-stakes decision-making. This Skill provides a structured taxonomy to identify these specific failure patterns in model outputs.

How do I audit an LLM's reasoning path for errors in data science tasks?

Auditing an LLM reasoning path involves applying a structured failure-mode checklist to the model response to surface potential missteps. This post-hoc analysis records flagged issues for corrective action.

Can I use this taxonomy to build guardrails for reliable LLM reasoning?

Yes, the failure mode taxonomy enables practical validation by guiding the design of guardrails for reliable reasoning. It provides editorial guidelines and checklists to mitigate identified cognitive failures in engineering contexts.

What is the best way to evaluate LLM outputs for high-stakes product decisions?

Evaluating LLM outputs for high-stakes product decisions is best achieved by applying structured checklists and best practices from a failure mode taxonomy. This guides critique and verification to catch cognitive blind spots.

When should I not rely on standard LLM outputs for high-stakes reasoning?

Standard LLM outputs should not be relied upon for high-stakes reasoning without post-hoc analysis when cognitive blind spots pose risks. A structured taxonomy is needed to verify the model reasoning path and apply corrective actions.