ground-truth-evaluation

Verify and calibrate ground-truth mappings for Oak semantic search.

7|3|Updated Jul 28, 2025
One-click install
npx skills add https://github.com/oaknational/oak-open-curriculum-ecosystem --skill ground-truth-evaluation
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ground-truth-evaluation
Source: https://github.com/oaknational/oak-open-curriculum-ecosystem/tree/main/.cursor/skills/ground-truth-evaluation
Command: npx skills add https://github.com/oaknational/oak-open-curriculum-ecosystem --skill ground-truth-evaluation

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Evaluate and validate ground truths used by Oak semantic search to ensure alignment with expected results and reliable benchmarks.

Core Features & Use Cases

  • Establish a repeatable process to review, adjust, and validate LessonGroundTruth and query-relevance mappings.
  • Use during benchmark runs, MRR diagnostics, and COMMIT-based assessments across multiple subjects and datasets.
  • Provide governance for data-quality, versioning, and traceability of ground-truth definitions.

Quick Start

Review an Oak semantic-search ground-truth dataset and run COMMIT-based validation to identify gaps and update relevance mappings.

Frequently Asked Questions about ground-truth-evaluation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I validate semantic search ground-truth mappings for benchmarking?

Validate semantic search ground-truth mappings by applying a repeatable workflow that enforces ground-truth definitions, updates relevance scores, and verifies bulk-downloaded data to ensure accurate benchmarking results.

What is the best way to diagnose MRR issues using ground-truth data?

Diagnose MRR issues by running COMMIT-based validation against your query-relevance mappings to identify gaps, verify definitions, and update relevance scores for precise search evaluation.

How does COMMIT-based review work for ground-truth calibration?

COMMIT-based review works by evaluating and validating ground-truth datasets across multiple subjects, ensuring alignment with expected results and providing governance for data-quality and versioning during calibration.

Can I use bulk-downloaded data to verify query-relevance mappings?

Yes, you can use bulk-downloaded data to verify query-relevance mappings, identify gaps, and update relevance scores within a repeatable workflow that enforces ground-truth definitions.

Why do I need to calibrate ground truths for semantic search?

You need to calibrate ground truths to ensure alignment with expected results, maintain data-quality governance, and establish reliable benchmarks for accurate semantic search evaluation.

Does ground-truth evaluation support multiple subjects and datasets?

Yes, ground-truth evaluation supports multiple subjects and datasets, applying COMMIT-based assessments and MRR diagnostics to validate LessonGroundTruth and query-relevance mappings across diverse data.