nemo-curator

Automate GPU-accelerated data curation for LLM training datasets.

1.0k|117|Updated Feb 26, 2026
One-click install
npx skills add https://github.com/OpenLAIR/dr-claw --skill nemo-curator-openlair
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: nemo-curator
Source: https://github.com/OpenLAIR/dr-claw/tree/main/skills/data-processing/nemo-curator
Command: npx skills add https://github.com/OpenLAIR/dr-claw --skill nemo-curator-openlair

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Data scientists and MLOps teams spend countless hours preparing training data for LLMs. Nemo Curator automates end-to-end data curation, combining fast deduplication, quality filtering, PII redaction, and multi-modal processing on GPU.

Core Features & Use Cases

  • Multi-modal deduplication: exact, fuzzy, and semantic deduplication across text, image, video, and audio data.
  • 30+ quality filters: aggressive cleaning to improve dataset quality and reduce toxicity and noise.
  • PII redaction & NSFW detection: protects sensitive information and content integrity at scale.
  • Large-scale GPU acceleration: scales across RAPIDS and CUDA-enabled clusters for massive datasets.
  • Use cases: prepare web-scraped corpora, curate product/data catalogs, or clean proprietary datasets for model training.

Quick Start

Install Nemo Curator and run a quick sample on a small dataset to see automatic deduplication and quality filtering in action.

Frequently Asked Questions about nemo-curator

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I prepare large-scale web-scraped corpora for LLM training?

Automated data curation prepares large-scale web-scraped corpora for LLM training by applying deduplication, 30+ quality filters, and PII redaction to yield high-quality, clean datasets ready for modeling.

What is GPU-accelerated data curation and how does it handle multi-modal data?

GPU-accelerated data curation uses RAPIDS and CUDA-enabled clusters to process text, images, video, and audio data, executing fast exact, fuzzy, and semantic deduplication at scale.

Does GPU-accelerated data curation work for cleaning proprietary datasets and product catalogs?

Yes, GPU-accelerated data curation works for cleaning proprietary datasets and product catalogs, applying aggressive quality filtering and NSFW detection to improve dataset integrity and reduce noise.

How do I redact PII and detect NSFW content in massive training datasets?

You redact PII and detect NSFW content in massive training datasets by applying automated data curation pipelines that identify and mask sensitive information while flagging inappropriate content at scale.

What's the best way to perform semantic deduplication across text and image data?

The best way to perform semantic deduplication across text and image data is using GPU-accelerated data curation tools that apply multi-modal processing to efficiently identify and remove semantic duplicates.

Do I need RAPIDS and CUDA-enabled clusters to run multi-modal data curation?

You need RAPIDS and CUDA-enabled clusters to achieve large-scale GPU acceleration for multi-modal data curation, ensuring fast processing across massive datasets for LLM training preparation.