mt-improve-skill

Evaluate and iteratively refine AI prompts through empirical execution scenarios.

Updated May 15, 2026
One-click install
npx skills add https://github.com/t-miura-024/tools --skill mt-improve-skill
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: mt-improve-skill
Source: https://github.com/t-miura-024/tools/tree/main/chezmoi/dot_config/opencode/skills/mt-improve-skill
Command: npx skills add https://github.com/t-miura-024/tools --skill mt-improve-skill

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps reduce ambiguity in AI instructions by replacing subjective prompt reviews with repeatable empirical evaluation and iterative improvement.

Core Features & Use Cases

  • Prompt Quality Evaluation: Runs unbiased execution scenarios and compares agent outcomes against predefined requirements and metrics.
  • Iterative Instruction Refinement: Identifies unclear instructions, traces failure phases, and applies targeted improvements until progress converges.
  • Use Case: Improve a frequently used skill, slash command, or agent prompt when the generated results are inconsistent and the root cause may be unclear guidance.

Quick Start

Ask the mt-improve-skill to evaluate and iteratively improve this AI prompt using unbiased execution scenarios and measurable criteria.

Frequently Asked Questions about mt-improve-skill

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I empirically evaluate and improve inconsistent AI prompts?

To empirically improve AI prompts, you evaluate them through unbiased execution scenarios, compare outcomes against predefined metrics, and iteratively apply targeted revisions until reliability converges.

What is the best way to fix ambiguous agent instructions that reduce reliability?

Fixing ambiguous agent instructions requires structured failure analysis to trace execution errors, followed by controlled prompt revisions guided by measurable evaluation criteria to eliminate unclear guidance.

Can I use empirical testing to optimize slash commands and task prompts?

Yes, empirical testing applies to slash commands, task prompts, and skill definitions by running scenario-based evaluations to measure performance and iteratively refine instructions for consistent results.

How does iterative instruction refinement work for AI agents?

Iterative instruction refinement works by executing scenario-based tests, analyzing failure phases structurally, and applying targeted prompt improvements repeatedly until the agent's output satisfies predefined metrics.

Why does my AI prompt generate inconsistent results across different scenarios?

AI prompts generate inconsistent results when instructions contain ambiguity, requiring scenario-based evaluation and structured failure analysis to identify unclear guidance and apply controlled revisions for reliable outcomes.

When should I replace subjective prompt reviews with empirical evaluation?

Replace subjective prompt reviews with empirical evaluation when generated results become inconsistent and the root cause is unclear, requiring repeatable execution scenarios and metric tracking to achieve robust improvement.