What problem does it solve?
This skill provides methodology and tooling for AI/ML security assessment—scanning prompts for prompt injection signatures, scoring model inversion risk, data poisoning exposure, and agent tool abuse. It maps findings to MITRE ATLAS techniques and recommends guardrail controls. It clarifies its focus on AI/ML system security rather than general application security.
Core Features & Use Cases
- AI Threat Scanner Tool: Scans prompts for injection signatures, scores risk, and maps findings to MITRE ATLAS techniques.
- Prompt Injection Detection & Jailbreak Assessment: Identifies direct and indirect prompt injection, jailbreak phrases, and system prompt extraction attempts.
- Model Inversion & Data Poisoning Risk: Assesses inversion risk by access level and data poisoning risk across fine-tuning scopes.
- MITRE ATLAS Coverage & Guardrails: Provides ATLAS mappings and pragmatic guardrail design patterns to mitigate findings.
- Cross-References & Workflow Guidance: Helps integrate findings into threat detection, incident response, and cloud-security contexts.
Quick Start
Run the threat scanner on your seed prompts to detect prompt injection and jailbreak risks.