botlearn-assessment

Evaluate agent capabilities across 10 dimensions and generate scores and improvement guidance.

9|2|Updated Mar 2, 2026
One-click install
npx skills add https://github.com/botlearn-ai/botlearn-skills --skill botlearn-assessment
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: botlearn-assessment
Source: https://github.com/botlearn-ai/botlearn-skills/tree/main/skills/botlearn-assessment
Command: npx skills add https://github.com/botlearn-ai/botlearn-skills --skill botlearn-assessment

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

OpenClaw's automated self-evaluation framework exposes a complete, objective view of a skill's readiness by measuring performance across 10 capability dimensions, highlighting strengths and gaps to guide production deployment.

Core Features & Use Cases

  • Multi-Dimension Coverage: Evaluates Task Efficacy, Information Retrieval, Reasoning, Code & Automation, Creative Generation, Tool Orchestration, Memory & Context, Cost Efficiency, Reliability, and Safety.
  • Bias Mitigation: Applies a -5% correction to CoT self-judged scores and provides transparent justification for each criterion.
  • Actionable Reporting: Produces per-dimension scores, overall readiness, improvement recommendations, and supports auditing for security and reliability.
  • Use Case: Quarterly production-readiness reviews to validate an agent's deployment risk profile before going live.

Quick Start

Run the OpenClaw self-evaluation flow to generate per-dimension scores and an improvement plan, then review the recommendations.

Frequently Asked Questions about botlearn-assessment

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate an AI agent's production readiness across multiple capability dimensions?

Agent production readiness evaluation uses a 10-dimension self-assessment framework covering Task Efficacy, Reasoning, Code, Reliability, and Safety. This generates per-dimension scores, bias-corrected reports, and improvement guidance to validate deployment risk profiles.

What is bias mitigation in AI self-evaluation and how does it affect scoring?

Bias mitigation in AI self-evaluation corrects subjective overestimation by applying a -5% correction to Chain-of-Thought self-judged scores. This adjustment ensures transparent, objective capability assessment and produces adjusted scores with detailed justification for each evaluated criterion.

How do I run a quarterly capability assessment for an AI agent before deployment?

Running a quarterly capability assessment involves executing an automated self-evaluation flow across 10 dimensions like Information Retrieval and Tool Orchestration. The process outputs a machine-readable metadata payload containing raw and adjusted scores, overall readiness metrics, and recommended skill upgrades.

What dimensions should I measure to assess an agent's reasoning and safety capabilities?

Assessing reasoning and safety capabilities requires evaluating specific dimensions including Reasoning & Planning, Memory & Context, Reliability, and Safety. The framework highlights strengths and gaps across these areas to guide deployment decisions and generate targeted improvement recommendations.

Can I audit an AI agent's reliability and safety scores for security compliance?

Auditing AI agent reliability and safety scores is fully supported by the assessment framework. It produces actionable reporting with per-dimension scores and transparent justifications, enabling security compliance reviews and validating the agent's overall production readiness profile.

What is the best way to identify capability gaps in an AI agent before going live?

Identifying capability gaps before deployment is best achieved through a multi-dimension self-assessment measuring Task Efficacy, Cost Efficiency, and Creative Generation. The evaluation highlights specific weaknesses and outputs an improvement plan with recommended skill upgrades to close the gaps.