benchmark-due-diligence

Runs adversarial multi-agent due diligence on benchmarks to separate marketing bubble from replicable playbook.

1.4k|216|Updated Oct 22, 2025
One-click install
npx skills add https://github.com/daymade/claude-code-skills --skill benchmark-due-diligence
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: benchmark-due-diligence
Source: https://github.com/daymade/claude-code-skills/tree/main/daymade-financial/benchmark-due-diligence
Command: npx skills add https://github.com/daymade/claude-code-skills --skill benchmark-due-diligence

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

When a founder, KOL, company, or product looks suspiciously successful, it is hard to tell real achievement from inflated marketing claims, and even harder to know which parts of their playbook you can actually copy. This Skill runs an adversarial due-diligence teardown that grades every claim, busts the bubble, and maps the validated playbook onto your own real resources.

Core Features & Use Cases

  • Adversarial claim verification: Parallel collection and verification agents grade every claim on an L1-L4 evidence scale with verdicts from confirmed to debunked, actively hunting falsifying evidence for headline stats like user counts, funding, and #1 rankings.
  • Attribution breakdown: Splits validated success into product strength, market timing, personal-IP marketing, and operations, separating replicable method from luck and timing.
  • Commissioner resource mapping: Maps the benchmark's playbook onto your actual owned assets with a four-tag table (borrow-able, not-replicable, already-doing, bubble-don't-copy) plus concrete action lists, while keeping your private context out of external web searches.
  • Use Case: You envy a competitor who claims 0-to-1M users and want to know what is real and what you can steal. The Skill verifies the claim's attribution and magnitude, debunks the inflated parts, and tells you exactly which of their moves work with your resources.

Quick Start

Run benchmark due diligence on [founder/company/product name] and map their playbook onto my resources.

Frequently Asked Questions about benchmark-due-diligence

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I verify if a competitor's claimed success is real?

Run this Skill on the competitor: collection agents gather claims with source URLs, then adversarial verification agents grade each claim L1-L4 and hunt for falsifying evidence. The output is a bubble-busting table sorted by most-inflated claim first.

How to tell replicable strategy from luck when studying a successful founder?

The Skill's attribution phase weights product strength, market timing, personal-IP marketing, and operations as percentage ranges, then explicitly labels each factor replicable method or non-replicable luck and timing. Phase 4 maps only the replicable parts onto your resources.

What is the difference between benchmark-due-diligence and deep-research?

Deep-research builds a neutral, trustworthy briefing on a topic, while this Skill assumes claims are inflated until proven otherwise and ends in a decision-oriented mapping of what you can personally use. Prefer this Skill for debunking and playbook extraction.

Does the due diligence process protect my private business information?

Yes. The Skill enforces a two-channel injection split: public facts about the benchmark go to all agents, while your private commissioner context reaches only the final mapping agent, which performs no external web searches.

Why does this Skill forbid running with context fork?

It is an orchestrator that spawns parallel collection and verification agents and may invoke other skills. Subagents cannot spawn subagents or call skills, so forking the context would silently break the entire fan-out.

What are the limitations of automated claim verification?

Verification quality depends on available public evidence; claims with no independent corroboration are marked doubtful rather than resolved. The Skill explicitly forbids guessing to fill gaps and flags unverified items as open questions.