arabic-agent-eval

Benchmark Arabic tool-calling performance for language models using the aae CLI.

1|Updated Apr 19, 2026
One-click install
npx skills add https://github.com/jackquelinunpredictable827/mkhlab --skill arabic-agent-eval-jackquelinunpredictable827
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: arabic-agent-eval
Source: https://github.com/jackquelinunpredictable827/mkhlab/tree/main/hermes-skills/arabic-agent-eval
Command: npx skills add https://github.com/jackquelinunpredictable827/mkhlab --skill arabic-agent-eval-jackquelinunpredictable827

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Arabic-speaking AI teams need measurable benchmarks to assess how LLMs call tools in Arabic, ensuring accuracy, safety, and dialect sensitivity.

Core Features & Use Cases

  • Quick evaluation: run a fast, single-provider check to gauge tool invocation behavior.
  • Full benchmark: execute a comprehensive run across multiple providers and dialects to compare performance.
  • Dataset & insights: view aggregated metrics across six evaluation categories (simple call, parameter extraction, multi-step reasoning, dialect handling, tool choice, and error handling).

Quick Start

Run a quick evaluation against an OpenAI model using the aae quick openai command to get started.

Frequently Asked Questions about arabic-agent-eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark LLM tool-calling performance in Arabic dialects?

The benchmark evaluates six categories: simple call, parameter extraction, multi-step reasoning, dialect handling, tool choice, and error handling, measuring how accurately language models invoke tools across various Arabic dialects and providers.

Can I run a quick evaluation for tool invocation against a single provider?

Yes, the quick evaluation mode runs a fast, single-provider check to gauge tool invocation behavior for language models using the aae quick openai command in your Python environment.

What environment do I need to run the Arabic LLM benchmark?

You need a Python environment with the aae CLI installed to execute the Arabic tool-calling benchmark and evaluate language model performance across dialects and providers.

How does the tool-calling benchmark handle different Arabic dialects?

The benchmark includes a dedicated dialect handling category within its full evaluation mode, testing how accurately language models invoke tools across different Arabic dialect variations and providers.

What is the difference between quick and full evaluation modes for AI agents?

Quick evaluation runs a fast single-provider check, while full benchmark mode executes a comprehensive run across multiple providers and dialects to compare aggregated performance metrics across six categories.