pytdc

Access drug discovery datasets and benchmarks from Therapeutics Data Commons.

1|2|Updated Apr 29, 2026
One-click install
npx skills add https://github.com/fuzzy-dynamics/strings --skill pytdc-fuzzy-dynamics
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: pytdc
Source: https://github.com/fuzzy-dynamics/strings/tree/main/packages/skills/pytdc
Command: npx skills add https://github.com/fuzzy-dynamics/strings --skill pytdc-fuzzy-dynamics

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires tdc, numpy, pandas, scipy, and includes scripts (resource) and references (resource) components.

What problem does it solve?

PyTDC centralizes access to AI-ready drug discovery datasets and benchmarks from the Therapeutics Data Commons, accelerating research workflows.

Core Features & Use Cases

  • Unified access to a broad catalogue of datasets (ADME, Toxicity, DTI, HTS, QM, generation)
  • Benchmark-ready splits and evaluation utilities to streamline model benchmarking
  • References and tooling for loading, splitting, processing, and evaluating models

Quick Start

Load PyTDC data, apply a scaffold split for a chosen task, and evaluate results with the included evaluators.

Frequently Asked Questions about pytdc

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I access AI-ready drug discovery datasets for benchmarking?

Access AI-ready drug discovery datasets by loading single-prediction, multi-prediction, and generation tasks across ADME, toxicity, DTI, HTS, and QM domains from the Therapeutics Data Commons. It provides centralized benchmark-ready data for model evaluation.

What is the Therapeutics Data Commons and what tasks does it support?

The Therapeutics Data Commons is a centralized resource providing AI-ready drug discovery datasets and benchmarks. It supports single-prediction, multi-prediction, and generation tasks across ADME, toxicity, DTI, HTS, and QM domains to enable end-to-end model benchmarking.

How do I load and split therapeutics datasets for model training?

Load and split therapeutics datasets using the included scripts, such as applying a scaffold split for your chosen task. The utilities provide benchmark-ready splits and evaluation tools to streamline data processing and model benchmarking.

Can I use pandas and numpy for drug discovery data processing?

Yes, you can use pandas and numpy for drug discovery data processing as they are core dependencies. The framework integrates these libraries to handle data loading, splitting, and processing utilities across the dataset catalogue.

What evaluation utilities are available for drug discovery benchmarks?

Evaluation utilities for drug discovery benchmarks include model evaluators and benchmark-ready splits for single-prediction, multi-prediction, and generation tasks. These tools streamline the evaluation of results across ADME, toxicity, DTI, HTS, and QM domains.

Does this Skill include references for the datasets in the catalogue?

Yes, the Skill includes a dataset references catalog providing citations for the broad catalogue of drug discovery datasets. This ensures traceability for the AI-ready data loaded for ADME, toxicity, DTI, HTS, and QM benchmarking tasks.