What problem does it solve?
PyTDC removes the friction of finding, splitting, evaluating, and comparing therapeutic machine learning datasets so researchers can focus on model development instead of dataset plumbing.
Core Features & Use Cases
- Curated Therapeutics Datasets: Access standardized ADME, toxicity, DTI, DDI, PPI, generation, and retrosynthesis datasets from a single workflow.
- Benchmarking and Evaluation: Run reproducible benchmark group evaluations with proper multi-seed protocols and task-appropriate metrics.
- Molecular Utilities: Convert molecular formats, filter compounds, transform labels, and query common bioinformatics identifiers.
- Oracle-Guided Optimization: Score generated molecules with target-specific and physicochemical oracles for goal-directed design.
- Use Case: A researcher can load an ADMET dataset, apply a scaffold split, train a model, evaluate it across five seeds, and rank generated molecules against QED or GSK3B objectives.
Quick Start
Use the pytdc skill to load a therapeutic dataset, apply the recommended split strategy, and prepare it for model training and evaluation.