ai-debug

Diagnose AI feature failures by auditing symptoms to root causes with the 4D framework.

16|3|Updated Oct 23, 2025
One-click install
npx skills add https://github.com/breethomas/bette-think --skill ai-debug
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ai-debug
Source: https://github.com/breethomas/bette-think/tree/main/plugins/bette-think/skills/ai-debug
Command: npx skills add https://github.com/breethomas/bette-think --skill ai-debug

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps product and engineering teams diagnose why an AI feature is underperforming, hallucinating, or behaving inconsistently by guiding a structured investigation that traces symptoms back to root causes.

Core Features & Use Cases

  • 4D Context Audit: Systematically inspects D1 (Job definition), D2 (Context), D3 (Discovery), and D4 (Defense) to pinpoint gaps and vulnerabilities.
  • Symptom-to-Cause Mapping: Maps common symptoms like hallucinations, inconsistency, generic outputs, slow responses, high cost, and demo-vs-prod mismatch to likely root causes and focus areas.
  • Actionable Recommendations: Produces a prioritized fixes list, a recommended quick win, diagnostic questions for each dimension, and options to export findings into issue trackers.
  • Use Case: Ideal for a PM or engineer triaging a production chat or assistant feature that started hallucinating domain details after a recent deployment.

Quick Start

Ask the AI Debug skill to audit a feature by describing the observed symptom and the expected behavior to receive a prioritized set of fixes and a recommended quick win.

Frequently Asked Questions about ai-debug

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I diagnose AI hallucinations and inconsistent outputs in a production feature?

Diagnose AI hallucinations by auditing symptoms to root causes using the 4D framework, which assesses job definition, context fidelity, discovery reliability, and failure defenses to pinpoint vulnerabilities and produce prioritized fixes.

What is the best way to debug a chat assistant that started giving generic or slow responses after deployment?

Debug a chat assistant by mapping symptoms like generic outputs or slow responses to likely root causes across the 4D audit dimensions, yielding a prioritized fixes list, diagnostic questions, and a recommended quick win.

Why does my AI feature perform differently in demos compared to production environments?

Demo-versus-prod mismatches occur due to context gaps or discovery failures; auditing job definition and context fidelity maps these performance mismatches to root causes and generates targeted fixes to align behavior.

Can I use this 4D context audit to triage incidents for high-cost AI features?

Yes, the 4D context audit applies to incident triage for high-cost AI features by inspecting defense mechanisms and discovery reliability, ensuring you receive a prioritized fixes list and a quick win to reduce costs.

How do I fix AI failures fast during feature testing without deep investigation?

Fix AI failures fast by describing the observed symptom and expected behavior to receive a structured 4D audit, which immediately outputs a recommended quick win and prioritized fixes without requiring manual investigation.