nemo-curator

Curate multimodal LLM training data with GPU-accelerated deduplication and filtering.

Updated Apr 23, 2026
One-click install
npx skills add https://github.com/Rawgrowth-Consulting/rawclaw-agent --skill nemo-curator-rawgrowth-consulting
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: nemo-curator
Source: https://github.com/Rawgrowth-Consulting/rawclaw-agent/tree/main/optional-skills/mlops/nemo-curator
Command: npx skills add https://github.com/Rawgrowth-Consulting/rawclaw-agent --skill nemo-curator-rawgrowth-consulting

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Curates high-quality, multimodal training data for LLMs by accelerating GPU-based data processing, deduplication, and quality filtering.

Core Features & Use Cases

  • GPU-accelerated curation across text, image, video, and audio
  • Fuzzy, exact, and semantic deduplication
  • PII redaction and NSFW detection
  • Scales with RAPIDS for large-scale dataset preparation

Use cases include preparing large language model training data, cleaning web data, and deduplicating large corpora.

Quick Start

Launch the Nemo Curator pipeline to begin GPU-accelerated data curation on your training dataset.

Frequently Asked Questions about nemo-curator

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I deduplicate large multimodal datasets for LLM training?

To deduplicate large multimodal datasets for LLM training, you can apply fuzzy, exact, and semantic deduplication. This process scales across text, image, video, and audio formats using GPU-accelerated RAPIDS environments.

Can I redact PII and detect NSFW content during data curation?

Yes, you can redact PII and detect NSFW content during data curation. These quality filtering steps are integrated directly into the GPU-accelerated pipeline to clean web data and prepare training corpora.

Does GPU-accelerated data curation work for text, image, video, and audio?

GPU-accelerated data curation works across text, image, video, and audio formats. It leverages RAPIDS multi-GPU scaling to process large-scale, multimodal datasets efficiently for LLM preparation.

What is the best way to clean web data for large language model training?

The best way to clean web data for large language model training is using a GPU-accelerated curation workflow. This approach combines deduplication, PII redaction, NSFW detection, and quality filtering to ensure high-quality inputs.

Do I need RAPIDS to scale data curation across multiple GPUs?

Yes, you need RAPIDS to scale data curation across multiple GPUs. RAPIDS provides the necessary multi-GPU scaling framework to handle large-scale dataset preparation and accelerate the curation pipeline.