run-eval

Run retrieval and injection evaluation tests on synthetic and real data.

Updated Apr 10, 2026
One-click install
npx skills add https://github.com/emmahyde/memesis --skill run-eval-emmahyde
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: run-eval
Source: https://github.com/emmahyde/memesis/tree/main/skills/run-eval
Command: npx skills add https://github.com/emmahyde/memesis --skill run-eval-emmahyde

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill simplifies the process of running evaluation suites to measure the accuracy and performance of retrieval and injection mechanisms within memory-driven AI systems.

Core Features & Use Cases

  • Execution of Evaluation Tests: Automates the running of synthetic and real observation benchmarks to assess model retrieval quality.
  • Baseline Comparison and Metrics Capture: Facilitates capturing baseline performance metrics and comparing current results for regression detection.
  • Use Case: A user wants to verify if recent changes reduced retrieval accuracy by running comprehensive evaluation pipelines with minimal command input.

Quick Start

Ask the AI to run the evaluation suite to check current retrieval performance on your dataset.

Frequently Asked Questions about run-eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I run an evaluation suite to test AI model retrieval accuracy?

To run an evaluation suite for AI model retrieval accuracy, you prompt the AI to execute the evaluation pipeline on your dataset. It automates testing of retrieval and injection mechanisms to measure performance metrics.

What is regression comparison in AI memory systems?

Regression comparison in AI memory systems is the process of capturing baseline performance metrics and comparing them against current test results. It helps detect if recent changes have degraded model retrieval accuracy or injection quality.

Can I use synthetic data to assess retrieval quality in memory-driven AI?

Yes, you can use synthetic data to assess retrieval quality in memory-driven AI. The evaluation tests support both synthetic and real observation benchmarks to comprehensively measure model robustness and accuracy.

How do I check if recent code changes reduced retrieval accuracy?

To check if recent changes reduced retrieval accuracy, run the full evaluation pipeline with minimal command input. It compares current test results against captured baseline metrics to immediately identify any performance regressions.

Do I need specific dependencies to perform model evaluation tests?

No specific dependencies are required to perform model evaluation tests. The Skill operates independently to handle baseline measurement, regression comparison, and full pipeline execution for quality assurance in AI development.

What is the best way to measure injection quality in AI memory systems?

The best way to measure injection quality in AI memory systems is by running comprehensive evaluation pipelines. This approach automates synthetic and real observation benchmarks to ensure model robustness and verify retrieval accuracy.