nemo-curator

Accelerate LLM training dataset curation with GPU computing and RAPIDS.

Updated Jul 3, 2026
One-click install
npx skills add https://github.com/LynxLabVN/office-agent --skill nemo-curator-lynxlabvn
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: nemo-curator
Source: https://github.com/LynxLabVN/office-agent/tree/main/agent-core/optional-skills/mlops/nemo-curator
Command: npx skills add https://github.com/LynxLabVN/office-agent --skill nemo-curator-lynxlabvn

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires nemo-curator, cudf, dask, rapids, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill addresses the challenge of preparing high-quality datasets for Large Language Models (LLMs) by providing GPU-accelerated data curation tools.

Core Features & Use Cases

  • GPU Acceleration: Significantly speed up data curation processes with RAPIDS.
  • Multimodal Data Support: Curation for text, images, video, and audio.
  • Use Case: Use this Skill to curate and prepare data for LLM training, including cleaning web data and deduplicating large corpora.

Quick Start

Install nemo-curator and use it to process your data with a single command.

Frequently Asked Questions about nemo-curator

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I accelerate data curation for LLM training using GPU computing?

You can accelerate data curation for LLM training using GPU computing with RAPIDS, which significantly speeds up data processing tasks compared to CPU-based methods.

Can I curate multimodal data like images and audio for LLM training?

Yes, you can curate multimodal data for LLM training. The system supports curation for text, images, video, and audio to prepare diverse training datasets.

What is the best way to deduplicate large text corpora for language models?

The best way to deduplicate large text corpora is using GPU-accelerated data curation tools, which efficiently clean web data and remove duplicates from large-scale datasets.

Do I need Python and Dask to prepare high-quality datasets with RAPIDS?

Yes, you need Python and specific libraries including Dask, cuDF, and RAPIDS to use this data curation workflow for preparing high-quality LLM training datasets.

Does nemo-curator work with Nemo for GPU-accelerated data preparation?

Yes, nemo-curator integrates with Nemo and RAPIDS to provide GPU-accelerated data preparation, allowing you to process large datasets efficiently for LLM training.