nemo-curator

Accelerate multi-modal LLM training data curation with GPU-accelerated pipelines.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/wwwillott/jobnimbus --skill nemo-curator-wwwillott
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: nemo-curator
Source: https://github.com/wwwillott/jobnimbus/tree/main/optional-skills/mlops/nemo-curator
Command: npx skills add https://github.com/wwwillott/jobnimbus --skill nemo-curator-wwwillott

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

NeMo Curator accelerates and scales the creation of high-quality, multi-modal training data for LLMs, reducing manual curation effort and improving dataset fidelity.

Core Features & Use Cases

  • GPU-accelerated quality filtering, deduplication, PII redaction, NSFW detection, and multilingual content assessment across text, image, audio, and video modalities
  • Multi-stage pipelines combining exact, fuzzy, and semantic deduplication with classifier-based filtering for production-grade data
  • Scalable, distributed workflows designed to prep large datasets like RedPajama v2 or The Pile

Quick Start

Instruct NeMo Curator to assemble a high-quality, multi-modal training dataset by applying deduplication, quality filtering, PII redaction, NSFW checks, and multi-GPU scaling.

Frequently Asked Questions about nemo-curator

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I prepare large-scale multimodal training data for LLMs?

You prepare large-scale multimodal training data by applying GPU-accelerated pipelines for deduplication, quality filtering, and PII redaction across text, image, audio, and video modalities to yield high-fidelity datasets.

What is GPU-accelerated data curation for LLM training?

GPU-accelerated data curation uses distributed processing and frameworks like NVIDIA RAPIDS to optimize and scale the creation of high-quality training datasets, significantly reducing manual effort and processing time.

Does NeMo Curator support multilingual content assessment and NSFW detection?

Yes, NeMo Curator supports multilingual content assessment and NSFW detection as part of its modular, classifier-based filtering stages designed to ensure dataset quality and safety across modalities.

How do I perform exact and fuzzy deduplication on massive text datasets?

You perform exact and fuzzy deduplication on massive datasets by using multi-stage GPU-accelerated pipelines that combine semantic deduplication with classifier-based filtering for production-grade data curation.

Do I need NVIDIA GPUs to run distributed data curation workflows?

Yes, you need NVIDIA GPUs because the distributed data curation workflows rely on GPU-acceleration via NVIDIA RAPIDS and NeMo Curator frameworks to scale processing for large datasets like RedPajama v2 or The Pile.