arabench

Benchmark Arabic language models across eight standardized quality categories.

13|4|Updated Apr 7, 2026
One-click install
npx skills add https://github.com/Moshe-ship/hurmoz --skill arabench-moshe-ship
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: arabench
Source: https://github.com/Moshe-ship/hurmoz/tree/main/arabench
Command: npx skills add https://github.com/Moshe-ship/hurmoz --skill arabench-moshe-ship

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Arabic LLM benchmarking across eight quality categories provides a structured, quantitative assessment of Arabic model capabilities.

Core Features & Use Cases

  • Comprehensive evaluation across 8 quality dimensions including translation, grammar, dialect, diacritization, summarization, QA, generation, and culture.
  • Cross-model comparison enabling researchers to rank Arabic LLMs and identify strengths and weaknesses.
  • Reproducible workflows with standardized metrics and outputs suitable for reporting and benchmarking across teams.

Quick Start

Run arabench run to execute the full benchmark and generate results.

Frequently Asked Questions about arabench

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark Arabic LLMs across multiple quality dimensions?

You can benchmark Arabic LLMs by running a standardized evaluation across eight quality categories, including translation, grammar, dialect, diacritization, summarization, QA, generation, and culture, to quantify model performance.

What metrics are used for evaluating Arabic language model performance?

Evaluating Arabic language model performance involves standardized metrics applied across eight categories to score translation, grammar, dialect, diacritization, summarization, QA, generation, and culture, producing an overall rating for cross-model comparison.

How do I compare multiple Arabic LLMs to identify strengths and weaknesses?

Comparing multiple Arabic LLMs requires running a reproducible benchmark workflow that scores each model across eight quality dimensions, enabling researchers to rank performance and identify specific strengths and weaknesses in dialect or generation tasks.

Can I generate reproducible benchmarking reports for Arabic QA and translation models?

Yes, you can generate reproducible benchmarking reports for Arabic QA and translation models by executing the benchmark workflow, which provides standardized category scoring and an overall rating suitable for cross-team reporting.

Does Arabic LLM benchmarking support evaluation of dialect and diacritization?

Arabic LLM benchmarking supports evaluation of dialect and diacritization directly, treating them as distinct quality categories alongside translation, grammar, summarization, QA, generation, and culture within the standardized multi-metric assessment.