codex-promptfoo-agentic-eval

Run and interpret Promptfoo-based AFM agentic evaluation suites for QA.

324|17|Updated Aug 8, 2025
One-click install
npx skills add https://github.com/scouzi1966/maclocal-api --skill codex-promptfoo-agentic-eval
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: codex-promptfoo-agentic-eval
Source: https://github.com/scouzi1966/maclocal-api/tree/main/.codex/skills/codex-promptfoo-agentic-eval
Command: npx skills add https://github.com/scouzi1966/maclocal-api --skill codex-promptfoo-agentic-eval

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Run and interpret the Promptfoo-based AFM agentic evaluation suite to enable end-to-end functional QA and model-quality assessment for agentic workflows.

Core Features & Use Cases

  • Run, expand, and interpret the Promptfoo agentic evaluation suite for AFM validation and quality measurement.
  • Distinguish failure types and report provenance (afm_bug, model_quality, harness_bug) with explicit provenance tags (afm_internal, primary_source, public_benchmark_inspired, synthetic).
  • Manage multiple harness configurations (structured, structured-stress, toolcall, agentic, frameworks, opencode) and review results from prepared matrices and datasets.
  • Provide ready-to-use prompts, configs, and datasets, with references to test reports and failure classifications for analysis.

Quick Start

Start by running the harness with the agentic profile to begin evaluating AFM.

Frequently Asked Questions about codex-promptfoo-agentic-eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I run agentic evaluation tests with Promptfoo for structured output and tool calling?

You can run agentic evaluation tests by selecting a harness configuration like structured, toolcall, or agentic profiles. The suite automates Promptfoo execution for end-to-end functional QA and model-quality assessment across these agentic workflows.

How do I distinguish model quality issues from harness bugs when an agentic evaluation fails?

To distinguish model quality issues from harness bugs during agentic evaluation, the suite provides explicit provenance reporting. It classifies failures into clear categories: afm_bug, model_quality, and harness_bug for accurate analysis.

Can I evaluate streaming and concurrency behaviors in agentic workflows using Promptfoo?

Yes, you can evaluate streaming and concurrency behaviors. The agentic evaluation suite scope explicitly covers streaming, concurrency, structured-output, tool-calling, grammar, and guided-json for end-to-end functional QA.

What is the best way to interpret Promptfoo test reports for AFM validation?

The best way to interpret Promptfoo test reports is by referencing prepared datasets and failure classifications under Scripts/ and test-reports. This provides explicit provenance tags like afm_internal and primary_source for analysis.