model-tester

Generate six-dimension radar profiles for AI models via automated cross-model evaluation.

8|Updated Mar 18, 2026
One-click install
npx skills add https://github.com/TerryFYL/ai-research-army --skill model-tester
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: model-tester
Source: https://github.com/TerryFYL/ai-research-army/tree/main/skills/model-tester
Command: npx skills add https://github.com/TerryFYL/ai-research-army --skill model-tester

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires yaml, openai, anthropic, google-generativeai, and includes scripts (resource) and assets (resource) components.

What problem does it solve?

This Skill automates standardized cross-model evaluation by generating a six-dimension radar profile for AI models.

Core Features & Use Cases

  • Six-dimension radar framework (Dim1–Dim6) with 18 test cases, scored by an AI Judge to form a searchable model profile.
  • End-to-end workflow for registering models, running tests, aggregating results, and visualizing radar data.
  • Cross-validation and comprehensive testing workflows to compare models and inform deployment decisions.
  • Output includes a composite radar, detailed task results, and stored histories for auditing.

Quick Start

Register a model and run the standard six-dimension radar workflow to generate a radar profile and composite results.

Frequently Asked Questions about model-tester

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark AI models across different dimensions?

You can compare AI models using a six-dimension radar framework with 18 test cases scored by an AI Judge, which generates a composite radar visualization and detailed task results to inform deployment decisions.

What is cross-validation for large language models?

Cross-validation uses an AI Judge to score 18 test cases across six dimensions, producing a composite radar visualization and stored histories for auditing model performance.

Does this model testing workflow support OpenAI and Anthropic APIs?

It supports OpenAI, Anthropic, and Google Generative AI dependencies, enabling cross-model evaluation across these platforms to produce standardized radar profiles.

How do I generate a radar chart for AI evaluation?

The workflow automates radar chart generation by scoring 18 test cases across six dimensions with an AI Judge, then visualizing the aggregated data into a composite radar profile.

What is the best way to automate cross-model evaluation?

Automating cross-model evaluation through a coordinated workflow that registers models, runs 18 standardized test cases, and aggregates scores into a visual radar profile provides end-to-end standardized benchmarking.

What are the limitations of using a six-dimension radar framework for model testing?

The framework is limited to six specific dimensions with 18 predefined test cases, meaning it may not fully capture specialized capabilities outside of information synthesis, judgement, consistency, precision, writing, and rule adherence.