scientific-critical-thinking

Critique AI/ML research plans and manuscript claims for evidence gaps.

Updated May 27, 2026
One-click install
npx skills add https://github.com/rauffatali/my-research-copilot --skill scientific-critical-thinking-rauffatali
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: scientific-critical-thinking
Source: https://github.com/rauffatali/my-research-copilot/tree/main/.agents/skills/scientific-critical-thinking
Command: npx skills add https://github.com/rauffatali/my-research-copilot --skill scientific-critical-thinking-rauffatali

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill prevents you from turning plausible-sounding ideas into expensive experiments or strong manuscript claims without decisive evidence, fair baselines, and leakage-safe evaluation.

Core Features & Use Cases

  • Adversarial research judgment: critiques hypotheses, novelty risk, methodology, evaluation trustworthiness, and result interpretation before implementation.
  • Decision-driven outputs: forces a clear next action (e.g., proceed, revise plan, gather evidence, or stop) so critique doesn’t stall progress.
  • Targeted gatekeeping across the workflow: supports research direction changes, dataset/leakage checks, experiment decisiveness, and alignment between claims and evidence.

Quick Start

Use scientific-critical-thinking to critique the proposed experiment plan and end with a go/stop decision and the exact evidence or baselines you still need.

Frequently Asked Questions about scientific-critical-thinking

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I critique a research hypothesis before running experiments?

To critique a research hypothesis, you pressure-test it by evaluating novelty risk, checking methodology, and identifying evidence gaps. This process ensures your hypothesis has decisive evidence and fair baselines before committing to implementation.

What is data leakage detection in machine learning evaluation?

Data leakage detection in machine learning evaluation identifies when training data compromises test set integrity, causing inflated performance metrics. It involves assessing dataset construction and evaluation protocols to prevent evidence gaps and overbroad conclusions in your research plan.

How do I evaluate experiment decisiveness for manuscript claims?

To evaluate experiment decisiveness for manuscript claims, you verify metric alignment, specify required baselines and ablations, and enumerate failure modes. This ensures your results interpretation supports the claims without overbroad conclusions or missing evidence.

Can I assess novelty risk for AI research proposals?

Yes, you can assess novelty risk for AI research proposals by pressure-testing the planned contributions against existing literature. This identifies the strongest and weakest cases of your manuscript claims and prevents committing to experiments without decisive evidence.

What is the best way to pressure-test a computer vision architecture change?

The best way to pressure-test a computer vision architecture change is to apply an adversarial evaluation protocol that checks method soundness, specifies required ablations, and produces a concrete decision recommendation with a clear next step.

Why does my evaluation protocol produce overbroad conclusions?

Your evaluation protocol produces overbroad conclusions when it lacks fair baselines, contains dataset leakage risks, or has misaligned metrics. Pressure-testing the research plan identifies these evidence gaps and enumerates failure modes before finalizing manuscript claims.