What problem does it solve?
Curating large-scale, high-quality training data for large language models is traditionally slow, expensive, and error-prone when processed on CPU hardware, especially for multimodal datasets and petabyte-scale web corpora.
Core Features & Use Cases
- 16x Faster Deduplication: GPU-accelerated fuzzy and exact deduplication processes 8TB datasets in 7.5 hours versus 120 hours on CPU.
- Comprehensive Quality Filtering: 30+ heuristic and classifier-based filters to remove low-quality, toxic, or irrelevant content from training datasets.
- Multimodal Support: Curate text, image, video, and audio datasets with built-in PII redaction, NSFW detection, and modality-specific preprocessing tools.
- Use Case: ML teams training LLMs on web-scraped data like Common Crawl can use this skill to automate end-to-end data cleaning, deduplication, and safety filtering to produce production-ready training datasets at 89% lower cost than CPU-based alternatives.
Quick Start
Use the nemo-curator skill to process your raw LLM training data, applying quality filtering, deduplication, and PII redaction to produce a clean, high-quality curated dataset ready for model training.