self-eval

Evaluate AI work quality with two-axis scoring and devil's advocate reasoning.

Updated Apr 2, 2026
One-click install
npx skills add https://github.com/4lerman/text_evaluator --skill self-eval-4lerman
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: self-eval
Source: https://github.com/4lerman/text_evaluator/tree/main/.agents/skills/engineering-advanced-skills/self-eval
Command: npx skills add https://github.com/4lerman/text_evaluator --skill self-eval-4lerman

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

The skill helps users honestly evaluate the quality of their AI work by using a two-axis scoring system that encourages devil's advocate reasoning and prevents score inflation.

Core Features & Use Cases

  • Two-axis scoring: Rates task ambition and execution quality independently, then combines via a lookup matrix.
  • Devil's Advocate reasoning: Requires arguing for both higher and lower scores before finalizing the evaluation.
  • Score persistence: Appends scores to a history file for continuous assessment across sessions.
  • Anti-inflation detection: Flags score clustering and prompts for re-evaluation if necessary.

Quick Start

After completing a task or work session, type /self-eval followed by a brief description of the work.

Frequently Asked Questions about self-eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate AI work quality and prevent score inflation over time?

You can evaluate AI work quality using a two-axis scoring system that rates task ambition and execution independently, then persists scores to a history file to detect inflation and clustering across sessions.

What is two-axis scoring for AI task assessment?

Two-axis scoring rates task ambition and execution quality independently, then combines them via a lookup matrix to provide a balanced evaluation that resists score inflation.

How do I stop inflated self-evaluation scores in AI work sessions?

To stop score inflation, the evaluation requires mandatory devil's advocate reasoning, arguing for both higher and lower scores before finalizing, and appends results to a history file to flag clustering.

Can I track AI evaluation history across multiple sessions?

Yes, the evaluation process appends scores to a persistent history file, analyzing past results to detect score clustering and prompting for re-evaluation if inflation is suspected.

What's the best way to conduct a devil's advocate reasoning evaluation for AI tasks?

The best way is to require arguments for both higher and lower scores before finalizing an evaluation, combining this devil's advocate reasoning with a two-axis lookup matrix for objective scoring.

Do I need external tools to run a text-based AI work evaluation?

No, the evaluation operates entirely in a text-based interface without external tools, relying solely on its internal scripts to manage scoring, devil's advocate reasoning, and history persistence.