pytdc

Access AI-ready drug discovery datasets and benchmarks for machine learning in therapeutics.

1|Updated Jan 14, 2026
One-click install
npx skills add https://github.com/Sologa/codex-pipeline --skill pytdc-sologa
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: pytdc
Source: https://github.com/Sologa/codex-pipeline/tree/main/.codex/skills/pytdc
Command: npx skills add https://github.com/Sologa/codex-pipeline --skill pytdc-sologa

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides access to curated datasets and benchmarks for drug discovery and development, streamlining machine learning model development for therapeutic applications.

Core Features & Use Cases

  • Access Diverse Datasets: Utilize datasets for molecular property prediction (ADME, toxicity), drug-target interactions, molecular generation, and more.
  • Standardized Benchmarks: Evaluate models using consistent metrics and data splits.
  • Use Case: A researcher wants to build a model to predict drug toxicity. They can use this Skill to load a relevant toxicity dataset, split it appropriately, train their model, and evaluate its performance using standard metrics.

Quick Start

Use the pytdc skill to load the 'Caco2_Wang' dataset for ADME prediction.

Frequently Asked Questions about pytdc

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I find AI-ready datasets for machine learning in drug discovery?

AI-ready drug discovery datasets provide curated molecular property prediction, drug-target interaction, and molecular generation benchmarks. They streamline therapeutic machine learning model development by supplying standardized data splits and evaluation metrics for consistent testing.

What standardized benchmarks are available for molecular property prediction?

Standardized benchmarks for molecular property prediction include ADME and toxicity datasets. They provide consistent evaluation metrics and standardized data splits, allowing researchers to train models and evaluate drug toxicity or absorption performance accurately.

How do I load a dataset for drug-target interaction analysis?

To load a dataset for drug-target interaction analysis, you access curated therapeutic machine learning benchmarks. These provide standardized data splits and chemical data processing utilities to facilitate training and evaluating interaction prediction models.

Can I use these datasets for de novo molecular generation tasks?

Yes, these datasets support de novo molecular generation tasks. They provide curated chemical data processing utilities and standardized benchmarks to evaluate generative models within therapeutic machine learning workflows.

Do I need external data processing libraries to use these therapeutic ML benchmarks?

No, you do not need external libraries for initial processing because these benchmarks include built-in chemical data processing utilities. They provide standardized data splits and evaluation metrics directly for therapeutic machine learning tasks.

What evaluation metrics are provided for drug toxicity prediction models?

Drug toxicity prediction models are evaluated using standardized metrics provided with the dataset. These benchmarks ensure consistent model evaluation by supplying predefined data splits and processing utilities for therapeutic machine learning.