Anti-Benchmark

Identify and challenge benchmark assumptions to improve evaluation rigor.

393|34|Updated Feb 10, 2026
One-click install
npx skills add https://github.com/Pthahnix/De-Anthropocentric-Research-Engine --skill anti-benchmark
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: Anti-Benchmark
Source: https://github.com/Pthahnix/De-Anthropocentric-Research-Engine/tree/main/skills/sop/anti-benchmark
Command: npx skills add https://github.com/Pthahnix/De-Anthropocentric-Research-Engine --skill anti-benchmark

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Identify and challenge benchmark assumptions to improve evaluation rigor.

Core Features & Use Cases

  • Ability to critique foundational benchmark claims and surface gaps
  • Works across domains where benchmarks govern evaluation results
  • Facilitates generating alternative benchmarking perspectives for robust testing

Quick Start

Provide a benchmark and its context to receive a structured critique and proposed alternatives.

Frequently Asked Questions about Anti-Benchmark

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I challenge benchmark assumptions to improve evaluation rigor?

To challenge benchmark assumptions, provide a benchmark and its context to receive a structured critique containing challenged assumptions, variants, and proposed alternatives for robust testing.

What is benchmark assumption critique and when do I need it?

Benchmark assumption critique is the process of stress-testing foundational claims in scientific and engineering benchmarks. You need it to surface gaps and improve evaluation rigor across domains where benchmarks govern results.

How do I generate alternative benchmarking perspectives for robust testing?

Generate alternative benchmarking perspectives by invoking a critique workflow on your benchmark and context, which returns an AntiBenchmarkResult containing challenged assumptions, variants, and proposed alternatives.

Does benchmark critique work across different scientific and engineering domains?

Yes, benchmark critique works across scientific and engineering domains where foundational claims must be stress-tested. It systematically surfaces gaps to improve evaluation rigor regardless of the specific benchmarking domain.

What do I need to provide to start a benchmark assumption critique?

You need to provide a benchmark and its context to the critique workflow. The system processes these inputs and returns a structured result with challenged assumptions, variants, and proposed alternatives.

What is the best way to surface gaps in foundational benchmark claims?

The best way to surface gaps in foundational claims is systematic assumption critique. By challenging benchmark assumptions, you receive a structured AntiBenchmarkResult exposing variants and proposed alternatives for more robust evaluation.