base-model-selector

Evaluate foundational LLM candidates against a custom rubric for fine-tuning projects.

4|1|Updated Dec 28, 2025
One-click install
npx skills add https://github.com/marcgreen/therapy-coach-finetune --skill base-model-selector
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: base-model-selector
Source: https://github.com/marcgreen/therapy-coach-finetune/tree/main/skills/base-model-selector
Command: npx skills add https://github.com/marcgreen/therapy-coach-finetune --skill base-model-selector

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill prevents wasted resources by ensuring you select the most appropriate base Large Language Model (LLM) for your fine-tuning project before you begin the costly process of data generation and training.

Core Features & Use Cases

  • Systematic Model Research: Guides an exhaustive search for candidate LLMs based on size, architecture, and domain fit.
  • Baseline Evaluation: Establishes a performance benchmark for potential base models against your specific domain rubric.
  • Informed Decision Making: Provides clear criteria for selecting a primary and backup model, or deciding if fine-tuning is even necessary.
  • Use Case: Before fine-tuning a model for therapeutic coaching, use this Skill to research and evaluate models like Qwen, Llama, and Mistral, determining which one offers the best starting point for conversational quality and empathy.

Quick Start

Use the base-model-selector skill to research and evaluate LLM candidates for a new fine-tuning project.

Frequently Asked Questions about base-model-selector

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I choose the right base LLM for fine-tuning before starting data generation?

To choose a base LLM for fine-tuning, systematically research candidate models and evaluate baseline performance against your custom domain rubric to determine project viability and potential gains.

What factors should I compare when selecting a foundational LLM for my project?

When selecting a foundational LLM, compare factors including the model's license, context length, architecture, domain-specific fit, and community support to ensure it meets your baseline evaluation criteria.

How do I evaluate if an LLM is a good fit for a specific domain like therapeutic coaching?

Evaluate an LLM for a specific domain by establishing a performance benchmark against a custom rubric, assessing conversational quality and empathy to determine if the baseline meets project requirements.

Do I need to perform baseline evaluation before fine-tuning a large language model?

Baseline evaluation is required before fine-tuning to prevent wasted resources, establishing a performance benchmark to decide if training is necessary or which candidate model offers the best starting point.

What is the best way to research candidate LLMs for a fine-tuning project?

The best way to research candidate LLMs is to conduct an exhaustive search based on model size, architecture, and domain fit, comparing quantitative evaluation metrics to select primary and backup models.