evaluate

Generate evaluation specs and run automated checks across artifact types.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/derickLeeT/visualize-plus --skill evaluate-derickleet
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluate
Source: https://github.com/derickLeeT/visualize-plus/tree/main/eval
Command: npx skills add https://github.com/derickLeeT/visualize-plus --skill evaluate-derickleet

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

AI-generated artifacts require structured, repeatable evaluation to ensure reliability, consistency, and actionable insights. This Skill provides a practical framework to assess quality across artifact types and generate a shareable, visual report.

Core Features & Use Cases

  • Structured evaluation workflow: generate evaluation specs, execute checks, and compile results into a visual report.
  • Artifact scope: evaluates HTML visualizations, code projects, documents, agent conversations, slide decks, dashboards, and more.
  • Use Case: a product team audits a new AI-generated dashboard and compares it against a benchmark to determine ship readiness.

Quick Start

Run the evaluation workflow on your artifact and review the generated HTML report.

Frequently Asked Questions about evaluate

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate the quality of AI-generated artifacts and visualizations?

To evaluate AI-generated artifacts, generate tailored evaluation specs, execute automated checks and manual scoring across defined dimensions, and compile a structured report with actionable evidence and recommendations.

What is the best way to assess if an AI-generated dashboard is ready to ship?

The best way to assess dashboard ship readiness is running a structured evaluation workflow that compares the AI-generated dashboard against a benchmark to determine quality and consistency.

Can I use a structured evaluation workflow for documents and code projects?

Yes, you can use a structured evaluation workflow for documents, code projects, slide decks, and agent conversations. It identifies the artifact type to generate tailored evaluation specs for each.

How does an automated evaluation report present quality assessment findings?

An automated evaluation report presents quality assessment findings by compiling manual scoring and automated checks into a visual HTML format, providing structured evidence and recommendations.

What do I need to generate a structured evaluation report for AI outputs?

To generate a structured evaluation report for AI outputs, you need a completed artifact like an HTML visualization or document. The workflow automatically identifies the type and runs tailored checks.