mteb-leaderboard

Retrieve and compare top model standings from MTEB and embedding benchmark leaderboards.

134|21|Updated Nov 12, 2025
One-click install
npx skills add https://github.com/letta-ai/skills --skill mteb-leaderboard
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: mteb-leaderboard
Source: https://github.com/letta-ai/skills/tree/main/ai/benchmarks/letta/terminal-bench-2/trajectory-only/mteb-leaderboard
Command: npx skills add https://github.com/letta-ai/skills --skill mteb-leaderboard

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill guides querying machine learning model leaderboards (MTEB, embedding benchmarks) and comparing standings across benchmarks with awareness of temporal validity.

Core Features & Use Cases

  • Identify authoritative sources and verify timestamps
  • Retrieve current leaderboard data and compare models
  • Document discrepancies and data recency for reproducible results

Quick Start

Retrieve the top models for a given benchmark and date, then summarize their performance and ranking.

Frequently Asked Questions about mteb-leaderboard

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I find the top-performing embedding models on MTEB benchmarks?

Query the MTEB leaderboard to retrieve current top-performing embedding models ranked by benchmark score. The Skill identifies authoritative sources, verifies timestamps, and returns live rankings with performance metrics and model standings.

Can I compare embedding model performance across multiple benchmarks?

Yes. Compare model standings across different embedding benchmarks by retrieving and cross-referencing leaderboard data from multiple sources. The Skill handles temporal alignment and data recency to ensure valid comparisons.

How do I verify the freshness and reliability of benchmark data?

The Skill documents data timestamps, verifies authoritative sources, and checks temporal validity. This ensures reproducible results and confirms that retrieved benchmark standings meet current eligibility criteria.

What's the best way to track embedding model improvements over time?

Retrieve leaderboard snapshots at different dates and compare model rankings across benchmarks. The Skill supports temporal awareness and provenance verification to track performance trends and identify which models have improved.

Does this work with HuggingFace embedding models?

Yes. The Skill queries MTEB and related embedding benchmarks where HuggingFace models are evaluated. Retrieve current standings and performance comparisons for models hosted on or benchmarked through HuggingFace.