ai-scoring

Score work against defined rubrics with per-criterion and overall scores.

11|1|Updated Feb 8, 2026
One-click install
npx skills add https://github.com/lebsral/DSPy-Programming-not-prompting-LMs-skills --skill ai-scoring
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ai-scoring
Source: https://github.com/lebsral/DSPy-Programming-not-prompting-LMs-skills/tree/main/skills/ai-scoring
Command: npx skills add https://github.com/lebsral/DSPy-Programming-not-prompting-LMs-skills --skill ai-scoring

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Score, grade, or evaluate work using AI against defined criteria, enabling consistent and justified assessments across essays, code reviews, support interactions, audits, and compliance checks.

Core Features & Use Cases

  • Define a rubric with 3-7 criteria and a clear scale.
  • Score each criterion independently with justification.
  • Calibrate scores using anchor examples to improve consistency.
  • Support multi-rater ensembles and guardrails for high-stakes scoring.

Quick Start

Provide a rubric and a sample submission to generate per-criterion scores and an overall score.

Frequently Asked Questions about ai-scoring

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I score essays and code reviews using an AI rubric?

Scoring essays and code reviews with an AI rubric involves defining 3-7 criteria with a clear scale, then evaluating each criterion independently to generate per-criterion scores and an overall score with justified explanations.

What is AI-driven rubric calibration and how do anchor examples improve consistency?

AI-driven rubric calibration uses anchor examples to standardize evaluation thresholds and improve scoring consistency. By mapping sample submissions to specific rubric scores, the AI aligns its assessments to established benchmarks.

Can I use multi-rater ensembles for high-stakes compliance audits?

Multi-rater ensembles support high-stakes compliance audits by aggregating multiple independent AI evaluations. This approach applies guardrails to per-criterion scores, reducing variance and ensuring justified, reliable audit results.

What's the best way to evaluate support interactions against defined criteria?

Evaluating support interactions against defined criteria requires a structured rubric with a clear scale. The AI scores each criterion independently, providing numeric scores with justified explanations for an overall quality assessment.

Does AI scoring work without providing a predefined rubric?

AI scoring requires a predefined rubric to function accurately. You must define 3-7 criteria and a clear scale upfront so the evaluation can generate per-criterion scores and deliver justified, consistent assessments.

Why does AI evaluation sometimes produce inconsistent scores for similar submissions?

Inconsistent AI evaluation scores often occur when anchor examples are not used for calibration. Providing anchor examples during rubric definition standardizes the scoring scale and enforces guardrails to reduce variance across similar submissions.