harness-eval

Evaluate software engineering tasks with a 15-task benchmark and scoring rubric.

13|6|Updated Apr 14, 2026
One-click install
npx skills add https://github.com/baekenough/second-brain --skill harness-eval
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: harness-eval
Source: https://github.com/baekenough/second-brain/tree/main/.claude/skills/harness-eval
Command: npx skills add https://github.com/baekenough/second-brain --skill harness-eval

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Structured SE task evaluation uses a 15-task benchmark to provide objective scores for agent performance.

Core Features & Use Cases

  • 15-task benchmark suite with a standardized scoring rubric.
  • Presets (all / quick) for rapid benchmarking and comparison.
  • Integrates with the evaluator-optimizer workflow to produce per-task scores and aggregate grade.

Quick Start

Run harness-eval with all benchmarks to generate the full results report.

Frequently Asked Questions about harness-eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
What is a structured software engineering benchmark and how does it evaluate tasks?

A structured software engineering benchmark evaluates agent performance using a standardized 15-task harness suite. It applies a standardized rubric across API design, data modeling, authentication flow, logging, and configuration to generate objective per-task scores and an aggregate grade.

How do I run a software engineering benchmark evaluation on all tasks?

To run a software engineering benchmark evaluation, use the all preset to execute the full 15-task harness suite. This generates per-task scores, an aggregate grade, and exports the formatted results to the predefined .claude/outputs path.

Can I use a quick preset for rapid software engineering benchmarking and comparison?

Yes, you can use the quick preset for rapid software engineering benchmarking. It runs a faster subset of the 15-task suite to produce objective scores quickly, enabling rapid comparison of agent performance across standard tasks.

Does the benchmark evaluation integrate with an evaluator-optimizer workflow?

Yes, the benchmark evaluation integrates directly with the evaluator-optimizer workflow. It formats per-task scores and aggregate grades into outputs optimized for the evaluator-optimizer pipeline to consume and process.

What software engineering tasks are covered by the benchmark scoring rubric?

The benchmark scoring rubric covers 15 tasks including API design, data modeling, authentication flow, logging, and configuration. It generates objective per-task scores and an aggregate grade for each.

Where are the software engineering benchmark results exported?

Software engineering benchmark results are exported to the predefined .claude/outputs path. The evaluation outputs are formatted specifically for consumption by the evaluator-optimizer pipeline.