promptfoo

Tests and evaluates LLM prompts across multiple models for regressions and performance.

2|Updated Jan 15, 2026
One-click install
npx skills add https://github.com/DTMC-marketplace/governance --skill promptfoo-dtmc-marketplace
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: promptfoo
Source: https://github.com/DTMC-marketplace/governance/tree/main/skills/promptfoo
Command: npx skills add https://github.com/DTMC-marketplace/governance --skill promptfoo-dtmc-marketplace

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill addresses the challenge of systematically testing and evaluating Large Language Model (LLM) prompts to ensure their effectiveness, consistency, and compliance.

Core Features & Use Cases

  • Systematic Prompt Testing: Evaluate prompts across multiple LLM providers and configurations.
  • Regression Detection: Identify unintended changes in prompt performance over time.
  • Performance Benchmarking: Compare different prompts or model versions to select the best performing ones.
  • Compliance Assessment: Aid in evaluating AI systems against regulatory requirements like the EU AI Act's Article 15.

Quick Start

Use the promptfoo skill to test and evaluate LLM prompts systematically.

Frequently Asked Questions about promptfoo

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I test and evaluate LLM prompts systematically?

To test and evaluate LLM prompts systematically, you can run them across multiple LLM providers and configurations to detect regressions and benchmark performance over time. This ensures prompt effectiveness and consistency.

What is LLM prompt regression detection?

LLM prompt regression detection is the process of identifying unintended changes in prompt performance over time. It helps ensure that updates to models or prompts do not degrade existing output quality and consistency.

How do I benchmark LLM prompts across different models?

To benchmark LLM prompts across different models, you compare various prompts or model versions against each other. This performance benchmarking allows you to select the best performing configurations for your specific use case.

Can I assess AI compliance against the EU AI Act Article 15?

Yes, you can assess AI compliance against EU AI Act Article 15 requirements by evaluating AI systems, implementing necessary controls, and systematically documenting your compliance findings and performance metrics.

Does prompt testing work with multiple LLM providers?

Prompt testing works with multiple LLM providers and configurations, allowing you to systematically evaluate prompt behavior. This cross-provider testing helps verify that prompts perform consistently regardless of the underlying model.