ai-evaluation-verification

Validate AI outputs with logic, first-principles, and reality gates.

Updated Jan 14, 2026
One-click install
npx skills add https://github.com/leobessa/claude-plugins-ai-fluency --skill ai-evaluation-verification
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ai-evaluation-verification
Source: https://github.com/leobessa/claude-plugins-ai-fluency/tree/main/skills/ai-evaluation-verification
Command: npx skills add https://github.com/leobessa/claude-plugins-ai-fluency --skill ai-evaluation-verification

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

AI outputs can mislead decisions; this skill provides a structured approach to validate reasoning, ensure alignment with first principles, and verify against external evidence to prevent confident nonsense and incomplete conclusions.

Core Features & Use Cases

  • Gate-based validation: logic check, first principles check, and reality check for every output.
  • Error-detection patterns: detect confident nonsense, false completeness, category errors, and anchoring on false premises.
  • Verification methods: spot-check, falsification testing, counter-prompting, source verification, and consistency testing.
  • Stopping rules and confidence scoring to document verification status and guide decision-making.

Quick Start

Activate this skill to apply three validation gates (logic, first principles, reality) to every AI output before acceptance.

Frequently Asked Questions about ai-evaluation-verification

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I validate AI outputs to prevent confident nonsense in decision-support workflows?

Spot-check AI reasoning by applying falsification testing, counter-prompting, and consistency testing to detect false completeness and category errors. These verification methods expose anchoring on false premises and temporal confusion, ensuring outputs maintain epistemic control.

What is gate-based validation for AI-generated content?

Falsification testing and counter-prompting verify AI reasoning by actively challenging outputs to expose anchoring on false premises and category errors. These methods, combined with source verification and consistency testing, ensure outputs withstand logical scrutiny before decisions.

Can I apply structured verification gates to AI code review and policy writing?

Verification status is documented using stopping rules and explicit confidence scoring. These mechanisms guide decision-making by providing a clear measure of how thoroughly an AI output has passed logic, first principles, and reality validation gates.

How do I detect false completeness and category errors in AI analysis?

Limitations of relying on AI outputs without verification include accepting confident nonsense, false completeness, and category errors as valid conclusions. Without gate-based validation, anchoring on false premises and temporal confusion can negatively impact decisions.