llm-evaluate

Fetch current pricing and specs to rank LLM models by cost and performance.

2|1|Updated Jan 27, 2026
One-click install
npx skills add https://github.com/lucidlabs-hq/agent-kit --skill llm-evaluate
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: llm-evaluate
Source: https://github.com/lucidlabs-hq/agent-kit/tree/main/.claude/skills/llm-evaluate
Command: npx skills add https://github.com/lucidlabs-hq/agent-kit --skill llm-evaluate

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Evaluates large language model options for cost and performance to help teams pick the right provider for a given use case.

Core Features & Use Cases

  • Pricing-aware evaluation across major providers (Anthropic, OpenAI, Google, DeepSeek, Mistral, xAI)
  • Performance and features scoring based on latency, context window, and capabilities relevant to chat, document analysis, or coding tasks
  • Use Case Mapping to align model choice with application needs such as chatbots, document processing, or code generation

Quick Start

Describe your use case and run the skill to receive prioritized model recommendations based on price, performance, and context.

Frequently Asked Questions about llm-evaluate

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I compare LLM pricing and performance across providers?

To compare LLM pricing and performance, you evaluate current model costs, specs, and latency across major providers. The system ranks models by transparently weighting cost, performance, context window, and features for your use case.

What is the best way to find a cost-effective LLM for document analysis?

Finding a cost-effective LLM for document analysis involves mapping use case requirements to model capabilities. The evaluation scores providers on context window, latency, and pricing to prioritize the best recommendations.

Does this LLM evaluation include latency and context window data?

Yes, the LLM evaluation includes latency and context window data. It fetches current model specs and benchmarks from providers like Anthropic, OpenAI, Google, DeepSeek, Mistral, and xAI to score performance.

Can I use this to evaluate models for code generation?

Yes, you can evaluate models for code generation. The evaluation aligns model choices with application needs by scoring capabilities relevant to coding tasks, chatbots, and document processing.

How do I get current LLM pricing for major providers?

To get current LLM pricing for major providers, the evaluation requires web access to fetch live cost data. It retrieves up-to-date pricing, benchmarks, and model specs to produce ranked recommendations.

When should I use a specialized provider comparison for LLM selection?

Use a specialized provider comparison when you need transparent scoring across cost, performance, context, and features. It is ideal when choosing between major providers like Anthropic, OpenAI, or Google for specific tasks.