partnership-research-eval

Grade partnership research artifacts against a four-dimension rubric and output verdict YAML.

1|2|Updated Apr 1, 2026
One-click install
npx skills add https://github.com/dimagi-internal/ace --skill partnership-research-eval
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: partnership-research-eval
Source: https://github.com/dimagi-internal/ace/tree/main/skills/partnership-research-eval
Command: npx skills add https://github.com/dimagi-internal/ace --skill partnership-research-eval

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Partnership research artifacts (deep web research reports and Connect/Dimagi capability-fit memos) often contain sourcing gaps, factual errors, or content misaligned with prospect expansion goals, leading to low-quality or risky prospect-facing materials. This skill provides consistent, objective LLM-as-judge quality grading to catch these issues before artifacts are used in partnership pitches.

Core Features & Use Cases

  • Four-dimension LLM-as-judge grading: Evaluates research artifacts across grounding, relevance, capability fit, and factual safety using a standardized rubric with hard deduction rules for critical issues like fabricated claims.
  • QA gating: Automatically skips evaluation if prior partnership-research-qa checks failed, avoiding wasted effort on invalid or incomplete artifacts.
  • Standardized verdict output: Writes a structured YAML verdict with dimension scores, weighted overall scores, and auto-surfaced severity concerns for aggregation into opportunity evaluation workflows. Use case: For a partnership development team building pitches for Connect opportunities, this skill automatically evaluates the quality of deep research and capability fit memos to catch unsourced claims, generic content, or factual errors before sharing materials with prospects.

Quick Start

Use the partnership-research-eval skill to evaluate the quality of the partnership research artifacts for the current prospect and generate a standardized quality verdict YAML.

Frequently Asked Questions about partnership-research-eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate the quality of partnership research artifacts?

LLM-as-judge quality evaluation uses a standardized rubric to grade research artifacts, applying hard deduction rules for fabricated claims. It assesses grounding, relevance, capability fit, and factual safety to catch unsourced claims before pitching.

Can I skip evaluation for research artifacts that failed QA checks?

Yes, QA gating automatically skips evaluation if prior partnership-research-qa checks failed. This avoids wasting effort on invalid or incomplete artifacts, ensuring only QA-passed research undergoes the four-dimension grading process.

How do I generate standardized verdicts for prospect evaluation pipelines?

Generating standardized verdicts involves producing a structured YAML file containing dimension scores, weighted overall scores, and auto-surfaced severity concerns. This output enables seamless aggregation into opportunity evaluation workflows for Connect opportunities.

What is the best way to catch factual errors in capability fit memos?

The best way to catch factual errors is applying a factual safety rubric with hard deduction rules for fabricated claims. This LLM-as-judge approach objectively grades memos for misalignment with prospect expansion theses before sharing with prospects.

What are the limitations of automated research artifact grading?

Automated research artifact grading is limited to evaluating artifacts only after QA gating confirms prior research checks passed. It requires deep web research reports and capability-fit memos as inputs and cannot evaluate incomplete or invalid artifacts.