huggingface-community-evals

Evaluate Hugging Face Hub models locally with inspect-ai and lighteval.

10.9k|724|Updated Nov 24, 2025
One-click install
npx skills add https://github.com/huggingface/skills --skill huggingface-community-evals-huggingface
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: huggingface-community-evals
Source: https://github.com/huggingface/skills/tree/main/skills/huggingface-community-evals
Command: npx skills add https://github.com/huggingface/skills --skill huggingface-community-evals-huggingface

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires inspect-ai, lighteval, vllm, transformers, accelerate, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides a local evaluation environment for Hugging Face Hub models, allowing users to assess model performance using inspect-ai and lighteval tools.

Core Features & Use Cases

  • Local Model Evaluation: Run evaluations on Hugging Face Hub models using inspect-ai and lighteval on local hardware.
  • Backend Selection: Choose between vLLM, Hugging Face Transformers, and accelerate for different backend needs.
  • Use Case: Ideal for backend selection, local GPU evaluations, and choosing between vLLM, Transformers, and accelerate for model performance assessment.

Quick Start

Run evaluations for the model 'meta-llama/Llama-3.2-1B' using inspect-ai with local GPU inference.

Frequently Asked Questions about huggingface-community-evals

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate Hugging Face Hub models locally on my own GPU?

You can evaluate Hugging Face Hub models locally by running inspect-ai and lighteval directly on your GPU. This Skill configures the local environment to assess model performance using your own hardware.

What is the best way to choose between vLLM, Transformers, and accelerate for local model evaluation?

The best way to choose between vLLM, Transformers, and accelerate is to use this Skill's backend selection feature. It allows you to run local GPU evaluations and compare model performance across these three distinct inference backends.

Can I use inspect-ai with lighteval for local GPU evaluations?

Yes, you can use inspect-ai with lighteval for local GPU evaluations. This Skill integrates both tools to provide a local evaluation environment for assessing Hugging Face Hub models on your own hardware.

Do I need vLLM to run local evaluations with Hugging Face Transformers?

You do not need vLLM to run local evaluations with Hugging Face Transformers. This Skill supports backend selection, allowing you to choose between vLLM, Hugging Face Transformers, and accelerate based on your specific evaluation needs.

Does this local model evaluation approach work with the meta-llama/Llama-3.2-1B model?

Yes, this local model evaluation approach works with models like meta-llama/Llama-3.2-1B. You can quickly start running inspect-ai evaluations with local GPU inference to assess its performance.