biomarker-discovery

Audit biomarker study designs for data leakage and validation rigor.

13|5|Updated May 4, 2026
One-click install
npx skills add https://github.com/awslabs/hcls-agent-skills --skill biomarker-discovery
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: biomarker-discovery
Source: https://github.com/awslabs/hcls-agent-skills/tree/main/skills/biomarker-discovery
Command: npx skills add https://github.com/awslabs/hcls-agent-skills --skill biomarker-discovery

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill addresses the high failure rate in biomarker development by enforcing rigorous methodological standards, preventing data leakage, and ensuring clinical utility is prioritized over mere statistical correlation.

Core Features & Use Cases

  • Methodological Auditing: Evaluates biomarker intent (prognostic vs. predictive vs. diagnostic) to ensure the study design matches the clinical question.
  • Leakage Prevention: Identifies and mitigates common pitfalls like temporal leakage, patient-level cross-validation errors, and improper feature selection.
  • Validation Framework: Guides the user through the hierarchy of validation, from internal nested cross-validation to external replication and prospective testing.

Quick Start

Use the biomarker-discovery skill to audit my proposed study design for a prognostic gene signature and identify potential data leakage risks.

Frequently Asked Questions about biomarker-discovery

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I prevent data leakage in high-dimensional omics biomarker discovery?

Preventing data leakage in biomarker discovery requires strict separation of feature selection and model training within nested cross-validation. This skill audits your study design to identify temporal leakage, patient-level cross-validation errors, and improper feature selection before analysis begins.

What's the best way to validate a prognostic gene signature for clinical utility?

Validating a prognostic gene signature requires progressing through a hierarchy from internal nested cross-validation to external replication and prospective testing. This skill enforces decision curve analysis and pre-specified thresholds to confirm diagnostic or prognostic utility beyond mere statistical correlation.

Why does my biomarker study design fail to match the clinical question?

Biomarker study designs fail when the intent—prognostic, predictive, or diagnostic—is not aligned with the clinical question. This skill provides methodological auditing to evaluate biomarker intent and ensure the study design matches the required clinical utility.

Can I use nested cross-validation for clinical and imaging data analysis?

Yes, nested cross-validation is essential for clinical and imaging data analysis to ensure methodological integrity. This skill applies a structured reasoning framework across high-dimensional omics, clinical, and imaging data to enforce rigorous validation standards.

When do I need decision curve analysis for biomarker validation?

Decision curve analysis is needed during biomarker validation to confirm that a model provides clinical utility over default strategies. This skill requires adherence to decision curve analysis alongside pre-specified thresholds to justify diagnostic or prognostic utility.