build-scoring

Design multi-dimensional evaluation rubrics with calibrated scales, thresholds, and adaptive weights.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/SujinHwang27/agent-os-lab --skill build-scoring
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: build-scoring
Source: https://github.com/SujinHwang27/agent-os-lab/tree/main/agent-os-factory-v2.0/.claude/skills/build-scoring
Command: npx skills add https://github.com/SujinHwang27/agent-os-lab --skill build-scoring

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Provides a structured, end-to-end approach to designing multi-dimensional evaluation rubrics with calibrated scales, thresholds, and function-adaptive weights to standardize decision-making across OS workflows.

Core Features & Use Cases

  • Scoring Context Design: Define 3-5 dimensions per context, with explicit scales and weights to reflect priorities.
  • Calibration & Examples: Provide concrete score descriptions and real-world examples for each score band to ensure consistent interpretation.
  • Thresholds & Guidance: Establish actionable stop/go/caution thresholds and recommended actions based on total scores.
  • Adaptive Weighting: Support context- and user-specific weight variants to adapt to different stages or personas.
  • Output & Governance: Produce a complete scoring framework and document it to domain-input/scoring_rubrics.md.

Quick Start

Create a complete scoring rubric for the current OS workflow and save it to domain-input/scoring_rubrics.md.

Frequently Asked Questions about build-scoring

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design multi-dimensional scoring rubrics for quality evaluation?

Multi-dimensional scoring rubrics are designed by defining 3-5 evaluation dimensions per context with explicit scales, calibrated score bands, and function-adaptive weights to standardize decision-making across OS workflows.

What is calibrated scoring and when do I need thresholds for go-no-go decisions?

Calibrated scoring uses concrete score descriptions with actionable stop, go, or caution thresholds to ensure consistent interpretation. You need it when establishing quality gates or triage processes across OS pipelines.

How do I set up adaptive weighting for different evaluation stages?

Adaptive weighting is set up by configuring context- and user-specific weight variants that adjust dimension priorities based on different workflow stages or personas within the evaluation rubric.

What's the best way to standardize triage and quality gate decisions?

The best way to standardize triage and quality gate decisions is creating a complete scoring framework with calibrated scales and explicit thresholds, then documenting it to domain-input/scoring_rubrics.md for governance.

Can I apply function-adaptive weights to agent OS performance scoring?

Yes, you can apply function-adaptive weights to agent OS performance scoring by defining dimensions with explicit scales and configuring adaptive safeguards to support context-specific weight variants.