nemo-evaluator-sdk

Evaluate large language models across 100+ benchmarks on Docker, Slurm, and cloud.

Updated Apr 11, 2026
One-click install
npx skills add https://github.com/hhhi21g/HealthCenter --skill nemo-evaluator-sdk-hhhi21g
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: nemo-evaluator-sdk
Source: https://github.com/hhhi21g/HealthCenter/tree/main/.codex/skills/nemo-evaluator
Command: npx skills add https://github.com/hhhi21g/HealthCenter --skill nemo-evaluator-sdk-hhhi21g

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires nemo-evaluator-launcher, docker, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill addresses the challenge of evaluating large language models (LLMs) across a wide range of benchmarks, providing a scalable solution for researchers and developers.

Core Features & Use Cases

  • Benchmarking: Offers access to over 100 benchmarks from various harnesses (MMLU, HumanEval, GSM8K, safety, VLM).
  • Multi-Backend Execution: Supports execution on local Docker, Slurm HPC, and cloud platforms.
  • Reproducibility: Ensures consistent benchmarking results with an enterprise-grade platform.
  • Use Case: Ideal for comparing the performance of different LLMs or for running comprehensive evaluations on a large dataset.

Quick Start

Run the nemo-evaluator-launcher to start evaluating a model using the following command:

nemo-evaluator-launcher run --config-dir . --config-name config

Frequently Asked Questions about nemo-evaluator-sdk

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark LLMs across multiple benchmarks with reproducible results?

Benchmark LLMs across 100+ benchmarks like MMLU and HumanEval using nemo-evaluator-sdk, which ensures consistent reproducible results through containerized execution and an enterprise-grade platform.

Can I run LLM evaluation workloads on Slurm HPC and cloud platforms?

Yes, LLM evaluation workloads support multi-backend execution on local Docker, Slurm HPC, and cloud platforms, allowing you to scale benchmarking tasks across various computing environments.

What benchmarks are supported for comparing large language models?

Supported benchmarks for comparing large language models include over 100 options across various harnesses such as MMLU, HumanEval, GSM8K, safety, and VLM evaluations.

Do I need Docker and nemo-evaluator-launcher to run LLM evaluations?

Yes, you need Docker and nemo-evaluator-launcher to run LLM evaluations, as they provide the required containerized execution environment and configuration launcher for the benchmarking process.

What's the best way to compare the performance of different LLMs on a large dataset?

The best way to compare different LLMs on large datasets is using a scalable benchmarking platform that evaluates models across diverse harnesses with reproducible results and multi-backend execution.