pytdc

Access curated drug discovery datasets and benchmarks for model training and evaluation.

1|Updated Mar 12, 2026
One-click install
npx skills add https://github.com/yf8578/clawomics --skill pytdc-yf8578
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: pytdc
Source: https://github.com/yf8578/clawomics/tree/main/skills/pytdc
Command: npx skills add https://github.com/yf8578/clawomics --skill pytdc-yf8578

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides access to a comprehensive suite of curated datasets and benchmarks for drug discovery and development, enabling rapid machine learning model development and evaluation.

Core Features & Use Cases

  • Access Diverse Datasets: Utilize over 200 datasets covering ADME, toxicity, drug-target interactions, molecular generation, and more.
  • Standardized Benchmarks: Evaluate models using predefined benchmark groups and splitting strategies (scaffold, cold-split).
  • Use Case: A researcher wants to build a model to predict drug toxicity. They can use this Skill to load multiple toxicity datasets, train their model using scaffold splits for robust generalization, and evaluate performance using standard metrics like ROC-AUC.

Quick Start

Load the Caco2_Wang ADME dataset and get its scaffold split.

Frequently Asked Questions about pytdc

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I get curated datasets for drug discovery machine learning?

You can access curated datasets for drug discovery machine learning by loading over 200 benchmarks covering ADME, toxicity, and drug-target interactions directly through this Skill. It provides standardized data splits and evaluation metrics for rapid model training and testing.

What benchmark datasets are available for predicting molecular toxicity and ADME?

Molecular toxicity and ADME prediction benchmarks are available as single-instance prediction datasets within the curated collection. You can load multiple toxicity datasets, train models using scaffold splits, and evaluate performance using standard metrics like ROC-AUC.

How do I use scaffold splits for robust generalization in cheminformatics models?

Scaffold splits for cheminformatics models are supported as predefined benchmark splitting strategies. You can load a dataset like Caco2_Wang and retrieve its scaffold split to ensure robust generalization during machine learning model training and evaluation.

Can I benchmark molecular generation models using standardized evaluation metrics?

Yes, you can benchmark molecular generation models using standardized evaluation metrics provided by this Skill. It facilitates machine learning model evaluation across molecular generation tasks alongside single-instance and multi-instance prediction benchmarks.

Does this Skill support drug-target interaction prediction datasets?

Drug-target interaction prediction datasets are fully supported as multi-instance prediction benchmarks. You can utilize these curated datasets for training and evaluating machine learning models focused on drug-target and drug-drug interactions.

What is the best way to evaluate drug discovery models across multiple datasets?

The best way to evaluate drug discovery models across multiple datasets is using predefined benchmark groups with standardized data splitting strategies. This Skill supports consistent evaluation across ADME, toxicity, and interaction datasets with built-in metrics.