jailbreak-taxonomy

Catalog jailbreak technique families and map defenses across AI safety systems.

4|Updated Apr 27, 2026
One-click install
npx skills add https://github.com/maruakshay/mii-ai-security --skill jailbreak-taxonomy
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: jailbreak-taxonomy
Source: https://github.com/maruakshay/mii-ai-security/tree/main/skills/jailbreak-taxonomy
Command: npx skills add https://github.com/maruakshay/mii-ai-security --skill jailbreak-taxonomy

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Jailbreak taxonomy helps security teams understand, categorize, and defend against a wide range of jailbreak mechanisms by providing a structured framework for detection and response.

Core Features & Use Cases

  • Technique cataloging across six mechanism families.
  • Defense mapping and escalation guidance for incident response.
  • Use Case: Red-team planning, policy hardening, and risk assessment across LLM deployments.

Quick Start

Review and map jailbreak families to defenses for your model, dataset, and workflow risk assessment.

Frequently Asked Questions about jailbreak-taxonomy

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
What is a jailbreak taxonomy and how does it map to LLM safety defenses?

A jailbreak taxonomy catalogs prompt injection and adversarial robustness mechanism families, mapping each to actionable defenses. It provides structured detection signatures and severity ratings for red-team testing, policy hardening, and risk assessment across LLM deployments.

How do I map prompt injection threats to defenses during red-team planning?

Map prompt injection threats to defenses by cataloging jailbreak techniques across six mechanism families. The taxonomy specifies technique details, detection signatures, and severity, enabling structured defense mapping and escalation guidance for security teams.

Can I use this taxonomy for system prompt hardening and risk assessment?

Yes, the taxonomy supports system prompt hardening and risk assessment across model prompts and outputs. It applies structured defense mapping to identify vulnerabilities and specify actionable defenses against adversarial robustness threats.

What are the main jailbreak technique families covered in an LLM threat model?

The taxonomy covers six main jailbreak mechanism families. It categorizes adversarial robustness techniques by specifying their technique details, detection signatures, and severity to assist in comprehensive threat modeling and response.

Does this jailbreak taxonomy provide detection signatures for incident response?

Yes, the taxonomy provides detection signatures and severity levels for each jailbreak mechanism family. This structured format directly supports incident response by offering defense mapping and escalation guidance for compromised AI safety systems.

What is the best way to structure threat model defenses against LLM jailbreaks?

Structure threat model defenses by mapping six jailbreak mechanism families to actionable defenses. The taxonomy outputs technique details, detection signatures, and severity ratings, ensuring comprehensive risk assessment and policy hardening.