self-eval

Evaluate completed AI work with calibrated two-axis ambition and execution scores.

Updated Apr 24, 2026
One-click install
npx skills add https://github.com/Veloxia-agency/VELOXIA-WEB --skill self-eval-veloxia-agency
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: self-eval
Source: https://github.com/Veloxia-agency/VELOXIA-WEB/tree/main/.claude/skills/engineering/skills/self-eval
Command: npx skills add https://github.com/Veloxia-agency/VELOXIA-WEB --skill self-eval-veloxia-agency

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps you avoid inflated self-assessments by replacing vague, single-number judgment with a calibrated evaluation of both task difficulty and execution quality.

Core Features & Use Cases

  • Two-axis scoring: Separates ambition from execution so the score reflects both what was attempted and how well it was done.
  • Devil's advocate review: Forces a skeptical pass that argues for lower and higher scores before finalizing the result.
  • Score persistence and trend checks: Records scores across sessions and flags repeated clustering that may signal bias.
  • Use cases: Ideal for post-task reflections, code review wrap-ups, debugging sessions, and any moment when you need an honest quality check on completed work.

Quick Start

Ask the assistant to evaluate the work you just completed and return a self-evaluation using the skill’s scoring rubric.

Frequently Asked Questions about self-eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I get an honest quality score for completed AI work instead of an inflated self-assessment?

To get an honest quality score for completed AI work, use a calibrated evaluation that separates task ambition from execution quality and applies a devil's advocate review to counter score inflation. This prevents vague single-number judgments by forcing skeptical analysis.

What is a two-axis rubric for AI task evaluation and when should I use it?

A two-axis rubric for AI task evaluation scores both the difficulty of what was attempted and how well the execution was performed. Use this calibrated scoring during post-task reflections, code review wrap-ups, and debugging session retrospectives for an unbiased assessment.

How do I run a devil's advocate review to check for score inflation in AI session retrospectives?

To run a devil's advocate review during AI session retrospectives, the evaluation process forces a skeptical pass that argues for both lower and higher scores before finalizing the result. This mandatory anti-inflation checking ensures the final calibrated score remains unbiased.

Can I track AI assessment quality scores across multiple sessions to detect bias?

Yes, you can track AI assessment quality scores across multiple sessions through score persistence. The system records calibrated scores over time and flags repeated clustering or trends, which helps identify potential bias in your retrospective evaluations.

Does this self-evaluation approach require any specific dependencies or external components?

No, this self-evaluation approach requires no external dependencies or components to function. You simply ask the assistant to evaluate your completed work and return a quality assessment using the built-in ambition-and-execution scoring rubric.

Why does my AI code review assessment need to separate ambition from execution?

Your AI code review assessment needs to separate ambition from execution because a single-number judgment often obscures whether a high score reflects an easy task done well or a complex task done poorly. A two-axis rubric provides a calibrated score reflecting both.