test-authority

Determine when to escalate from haiku/sonnet to sonnet/opus models.

326|35|Updated Nov 23, 2025
One-click install
npx skills add https://github.com/athola/claude-night-market --skill test-authority
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: test-authority
Source: https://github.com/athola/claude-night-market/tree/main/plugins/abstract/skills/escalation-governance/test-authority.md
Command: npx skills add https://github.com/athola/claude-night-market --skill test-authority

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Use when you need to test appropriate escalation when authority/context genuinely requires it.

Core Features & Use Cases

  • Realistic escalation scenario with decision points
  • Distinguish between knowledge gaps vs capability gaps
  • Documented criteria for escalation vs investigation

Quick Start

Execute a pressure test to determine escalation ethics and boundaries.

Frequently Asked Questions about test-authority

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
When should I escalate from a lower-capability model to a higher-capability model in agent workflows?

Escalate when the task presents genuine complexity or reasoning depth that exceeds current capability—novel patterns, high-stakes decisions, or knowledge gaps requiring deeper analysis. Escalation is justified when the problem itself demands it, not due to thrashing, time pressure, or unfounded safety concerns.

How do I distinguish between a knowledge gap and a capability gap when deciding to escalate?

A knowledge gap means the model lacks specific information but has sufficient reasoning capacity to process it if provided. A capability gap means the task requires reasoning depth, pattern recognition, or judgment beyond the current model's architecture. Escalate for capability gaps; investigate or provide context for knowledge gaps.

What criteria should guide escalation decisions in multi-tier model workflows?

Escalation decisions require a formal framework with documented criteria: define what constitutes genuine complexity, establish decision checkpoints where escalation is evaluated, exclude false triggers like performance pressure or unvalidated safety flags, and ensure all escalations are traceable and justified for audit and transparency.

Can I test escalation authority boundaries without affecting production agent behavior?

Yes. Use a pressure-test scenario with defined decision points to evaluate when escalation is warranted versus unnecessary. This lets you validate your escalation protocol, identify judgment gaps, and refine authority thresholds in a controlled environment before deploying to live workflows.

What makes an escalation decision traceable and justified?

Document the specific condition triggering escalation, the reasoning framework applied, which criteria were evaluated, and the outcome. Implement agent-schema guidance that captures decision metadata, ensuring every escalation has an audit trail linking the trigger, decision logic, and result for compliance and learning.

Related Skills