pytdc

Access AI-ready drug discovery datasets with standardized evaluation splits.

1|Updated Mar 4, 2026
One-click install
npx skills add https://github.com/Hung-3008/agusta --skill pytdc-hung-3008
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: pytdc
Source: https://github.com/Hung-3008/agusta/tree/main/.agents/skills/pytdc
Command: npx skills add https://github.com/Hung-3008/agusta --skill pytdc-hung-3008

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires tdc, numpy, pandas, and includes scripts (resource) and references (resource) components.

What problem does it solve?

PyTDC consolidates therapeutics data assets into an AI-ready, benchmarking-oriented platform, enabling researchers to access curated datasets, standardized splits, and molecular oracles for drug discovery tasks.

Core Features & Use Cases

  • Access ADME, Tox, DTI, and generation datasets with standardized evaluation metrics.
  • Use scaffold-based splits and other partition strategies for realistic model assessment; leverage molecular oracles and benchmarks for end-to-end workflows.
  • Run reproducible experiments with provided templates and references to example scripts and utilities.

Quick Start

Install PyTDC via pip and begin loading datasets for benchmarking.

Frequently Asked Questions about pytdc

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I access AI-ready drug discovery datasets for benchmarking?

Access AI-ready drug discovery datasets by loading curated therapeutics data assets directly through the platform. It provides standardized splits, molecular oracles, and ADME, Tox, and DTI datasets for reproducible benchmarking workflows.

What is a scaffold split and when do I need it for molecular model evaluation?

A scaffold split partitions molecular data based on chemical structure to ensure realistic model assessment. You need it for generation datasets and benchmarks to evaluate generalization on novel chemical scaffolds.

How do I run reproducible experiments for ADME and toxicity prediction?

Run reproducible experiments for ADME and toxicity prediction by using provided templates, example scripts, and standardized evaluation metrics. The platform loads structured guidance and reference utilities to support setup and benchmarking.

Does the therapeutics data commons platform work with pandas and numpy?

Yes, the therapeutics data commons platform works with pandas and numpy as core dependencies. It relies on these libraries to load, manipulate, and structure AI-ready datasets for drug discovery tasks.

Can I use molecular oracles for end-to-end drug discovery workflows?

Yes, you can use molecular oracles for end-to-end drug discovery workflows. The platform provides oracles alongside standardized splits and generation datasets to support comprehensive benchmarking from data loading to evaluation.

What types of drug discovery benchmarks are available for interaction datasets?

Available drug discovery benchmarks for interaction datasets include multi-instance interaction benchmarks and DTI datasets. These benchmarks provide standardized evaluation metrics and partition strategies for realistic model assessment.