model-escalation

Track task failures and recommend higher-tier model usage by threshold.

1|Updated May 21, 2026
One-click install
npx skills add https://github.com/hiddink-ai/hiddink-harness --skill model-escalation-hiddink-ai
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: model-escalation
Source: https://github.com/hiddink-ai/hiddink-harness/tree/main/templates/skills/model-escalation
Command: npx skills add https://github.com/hiddink-ai/hiddink-harness --skill model-escalation-hiddink-ai

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

This Skill addresses the inefficiency of using underpowered models for complex tasks by tracking failure rates and providing data-driven recommendations for model upgrades.

Core Features & Use Cases

  • Task Outcome Tracking: Monitors success and failure rates per agent and model type.
  • Escalation Advisories: Suggests model upgrades (e.g., Haiku to Sonnet to Opus) based on configurable failure thresholds.
  • Cost Awareness: Provides estimated cost multipliers to ensure performance improvements are balanced against budget constraints.

Quick Start

Enable the model escalation skill to monitor task outcomes and provide upgrade recommendations for the current agent session.

Frequently Asked Questions about model-escalation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I track task execution outcomes to recommend AI model upgrades?

Track task execution outcomes using shell-based outcome recording to analyze failure thresholds and recommend AI model upgrades. This monitors session-scoped performance metrics within agent workflows to identify bottlenecks and suggest higher-tier model usage for complex tasks.

Why does my AI agent repeatedly fail complex tasks and need model escalation?

AI agents fail complex tasks when using underpowered models, requiring model escalation based on tracked failure rates. Monitoring outcome data identifies performance bottlenecks and triggers data-driven recommendations to upgrade models like Haiku to Sonnet to Opus.

What is the best way to balance AI performance improvements against budget constraints?

Balance AI performance improvements against budget constraints by utilizing cost awareness features that provide estimated cost multipliers. This ensures suggested model upgrades for performance bottlenecks are evaluated against budget limits before escalating to higher-tier models.

Can I monitor agent success rates per model type without blocking execution?

You can monitor agent success rates per model type without blocking execution using shell-based outcome recording. This maintains session-scoped performance metrics asynchronously within agent workflows to track failures and suggest model upgrades.

When do I need to configure failure thresholds for model escalation advisories?

Configure failure thresholds for model escalation advisories when tracking task execution outcomes to identify underpowered models. Setting these thresholds triggers data-driven recommendations to upgrade models based on monitored failure rates within agent workflows.