langfuse-dataset-setup

Configure Langfuse datasets with evaluation metrics and judge prompts.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/mberto10/mberto-compound --skill langfuse-dataset-setup
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: langfuse-dataset-setup
Source: https://github.com/mberto10/mberto-compound/tree/main/.agents/skills/langfuse-dataset-setup
Command: npx skills add https://github.com/mberto10/mberto-compound --skill langfuse-dataset-setup

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill streamlines the process of setting up new datasets within Langfuse, including defining evaluation dimensions, configuring scoring, and preparing judge prompts for LLM-based or human review.

Core Features & Use Cases

  • Dataset Creation: Guides users through defining dataset purpose, evaluation dimensions, and method.
  • Evaluation Configuration: Provides templates and guidance for setting up score types and judge prompts.
  • Use Case: A product manager needs to set up a new dataset for evaluating chatbot responses for accuracy and helpfulness. This Skill helps them define the dataset, configure the scoring metrics, and generate the necessary prompts for an LLM judge.

Quick Start

Use the langfuse dataset setup skill to create a new dataset for evaluating response accuracy and helpfulness.

Frequently Asked Questions about langfuse-dataset-setup

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I set up Langfuse datasets for evaluating LLM responses?▼

To set up Langfuse datasets, you need to define the dataset purpose, evaluation dimensions, score configurations, and judge prompts. This process structures your data for both LLM-as-judge and human review workflows.

Can I configure LLM-as-judge evaluations for specific metrics like helpfulness?▼

Yes, you can configure LLM-as-judge evaluations by generating specific judge prompts using provided templates. The setup supports defining evaluation metrics such as accuracy, helpfulness, and relevance to score responses.

What is the best way to define evaluation dimensions for prompt engineering datasets?▼

The best way to define evaluation dimensions is to outline specific metrics like accuracy and relevance during dataset creation. Configuring score types and judge prompts ensures structured prompt engineering evaluation.

Do I need specific templates to generate judge prompts for human review?▼

Yes, generating judge prompts for human review utilizes specific templates provided during the evaluation configuration. These templates guide the setup of score types and prompt structures for your data management workflow.

Does Langfuse dataset setup support custom score configurations for chatbot accuracy?▼

Yes, Langfuse dataset setup supports custom score configurations for evaluating chatbot accuracy. It guides you through defining the dataset purpose and configuring the necessary scoring metrics for your evaluations.