nemo-curator

Automate AI training dataset curation and deduplication with GPU acceleration.

539|39|Updated May 1, 2026
One-click install
npx skills add https://github.com/Tommy-yw/RunbookHermes --skill nemo-curator-tommy-yw
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: nemo-curator
Source: https://github.com/Tommy-yw/RunbookHermes/tree/main/optional-skills/mlops/nemo-curator
Command: npx skills add https://github.com/Tommy-yw/RunbookHermes --skill nemo-curator-tommy-yw

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires nemo-curator, cudf, dask, rapids, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill addresses the challenge of preparing high-quality training datasets for large language models, by offering fast, accurate, and scalable data curation for text, image, video, and audio content.

Core Features & Use Cases

  • GPU Acceleration: Utilizes RAPIDS to accelerate data processing and deduplication tasks by 16×.
  • Data Curation: Provides features for deduplication, quality filtering, PII redaction, and NSFW detection.
  • Multi-modal Support: Works with text, images, video, and audio data types.
  • Use Case: Ideal for data scientists who need to prepare high-quality training datasets from diverse data sources like Common Crawl.

Quick Start

Use the nemo-curator skill to preprocess the Common Crawl dataset for training purposes.

Frequently Asked Questions about nemo-curator

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I prepare high-quality training data for large language models from Common Crawl?

To prepare high-quality training data for large language models, you can use nemo-curator to automate data curation, deduplication, and quality filtering on Common Crawl datasets. It applies PII redaction and NSFW detection, supporting text, image, video, and audio formats.

What's the best way to accelerate data deduplication and quality filtering for AI training datasets?

The best way to accelerate data deduplication and quality filtering is through GPU acceleration. This skill utilizes RAPIDS to accelerate data processing tasks by 16×, significantly reducing the time needed to curate large-scale AI training datasets.

Can I use GPU acceleration to redact PII and detect NSFW content in multi-modal datasets?

Yes, you can use GPU acceleration to redact PII and detect NSFW content across multi-modal datasets. The skill supports text, images, video, and audio data types, utilizing RAPIDS for fast and scalable quality control measures.

Do I need RAPIDS and Dask to perform large-scale data curation and deduplication?

Yes, you need RAPIDS and Dask dependencies to perform large-scale data curation and deduplication. The skill relies on these frameworks, alongside cudf, to execute GPU-accelerated processing and manage distributed data workflows efficiently.

What are the limitations of using GPU-accelerated data curation for diverse data sources?

The primary limitation of GPU-accelerated data curation is its dependency on specific frameworks like RAPIDS, cudf, and Dask. Users must ensure their environment supports these dependencies to successfully execute the 16× accelerated processing workflows for multi-modal data.

How does nemo-curator handle quality filtering for multi-modal AI training datasets?

nemo-curator handles quality filtering for multi-modal AI training datasets by automating deduplication, PII redaction, and NSFW detection. It processes text, images, video, and audio, using GPU acceleration to ensure accurate and scalable data preparation.