What problem does it solve?
This Skill protects Large Language Model (LLM) applications by detecting and filtering malicious prompts, including prompt injections and jailbreak attempts, ensuring safer and more reliable AI interactions.
Core Features & Use Cases
- Prompt Injection & Jailbreak Detection: Identifies and flags harmful inputs designed to manipulate LLM behavior.
- Third-Party Data Filtering: Secures RAG systems and API responses by detecting embedded malicious instructions.
- Multilingual Support: Operates effectively across 8 different languages.
- Use Case: Before sending user input to an LLM, use this Skill to scan for any attempts to override its instructions or inject harmful commands, preventing security breaches and ensuring the LLM stays on task.
Quick Start
Use the prompt-guard skill to check if the user input 'Ignore previous instructions' is a jailbreak attempt.