experiment-analyzer

Analyze LLM experiment results to produce summaries, comparisons, and actionable insights.

150|23|Updated Feb 3, 2026
One-click install
npx skills add https://github.com/datadog-labs/agent-skills --skill experiment-analyzer-datadog-labs
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: experiment-analyzer
Source: https://github.com/datadog-labs/agent-skills/tree/main/dd-llmo/experiment-analyzer
Command: npx skills add https://github.com/datadog-labs/agent-skills --skill experiment-analyzer-datadog-labs

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Analyze and interpret LLM experiment results, enabling quick decisions on which designs perform best.

Core Features & Use Cases

  • Compare single or multiple experiments to reveal performance differences, data quality, and reliability.
  • Generate actionable insights, summaries, and UI links to Datadog experiment dashboards.
  • Support Q&A and exploratory modes to answer specific questions or browse experiment results.

Quick Start

Run /experiment-analyzer with one or two experiment IDs and an optional question to begin analysis.

Frequently Asked Questions about experiment-analyzer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I analyze and compare LLM experiment results?

To analyze LLM experiment results, you can compare single or multiple experiments to evaluate performance differences, data quality, and reliability. This process generates concise summaries and actionable insights with direct Datadog UI dashboard links.

What is the best way to summarize LLM experiment performance?

Summarizing LLM experiment performance involves reviewing experiment events and metric values to produce concise structured insights. You can use exploratory mode to browse results or Q&A mode to answer specific performance questions.

Can I use Datadog APIs to compare two LLM experiments directly?

Yes, you can compare two LLM experiments directly by providing two experiment IDs. The analysis uses Datadog APIs to retrieve experiment summaries, events, and metric values, revealing performance differences and reliability between the two runs.

How do I get actionable insights from LLM experiment metrics?

Getting actionable insights from LLM experiment metrics requires analyzing dimension values and event data across experiments. The output provides structured recommendations and direct links to Datadog dashboards to help you decide which designs perform best.

Does analyzing LLM experiments require specific experiment IDs?

Analyzing LLM experiments requires providing one or two experiment IDs. You can also include an optional specific question to activate Q&A mode for targeted analysis instead of a general exploratory review of the experiment data.