What problem does it solve?
This Skill automates the complex and time-consuming process of preparing high-quality datasets for training Large Language Models (LLMs), significantly reducing the cost and effort involved.
Core Features & Use Cases
- GPU-Accelerated Processing: Leverages NVIDIA RAPIDS for massive speedups in deduplication, filtering, and PII redaction.
- Multimodal Support: Handles text, image, video, and audio data curation.
- Advanced Deduplication: Offers exact, fuzzy, and semantic deduplication to ensure data uniqueness.
- Quality & Safety: Includes over 30 heuristics for quality filtering, PII redaction, and NSFW detection.
- Use Case: Prepare a high-quality, deduplicated, and PII-redacted text dataset of 10TB from web scrapes for LLM training, achieving this in hours instead of weeks.
Quick Start
Use the nemo-curator skill to prepare a high-quality, deduplicated, and PII-redacted text dataset of 10TB from web scrapes for LLM training.