One-click install
npx skills add https://github.com/hoanghn61/.agents --skill ai-threat-testing-hoanghn61
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ai-threat-testing
Source: https://github.com/hoanghn61/.agents/tree/main/skills/ai-threat-testing
Command: npx skills add https://github.com/hoanghn61/.agents --skill ai-threat-testing-hoanghn61

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill helps security teams evaluate whether LLM applications are vulnerable to OWASP GenAI threats by running targeted offensive tests and producing evidence-based results.

Core Features & Use Cases

  • OWASP LLM Top 10 coverage: Runs specialized agent workflows for prompt injection, insecure output handling, training data poisoning, resource exhaustion, supply chain issues, excessive agency, model extraction, vector database poisoning, overreliance, and logging/monitoring bypass.
  • Evidence-driven exploitation: Captures prompts, outputs, logs, and execution artifacts to support reproducible proof-of-concept findings.
  • Pentest workflow integration: Combines AI-specific findings with traditional penetration testing for a unified security assessment and remediation roadmap.
  • Use Case: Assess a production chatbot and its RAG pipeline before a release by running a full OWASP Top 10 assessment to generate prioritized vulnerabilities with PoCs and hardening guidance.

Quick Start

Run an authorized full OWASP Top 10 assessment against your LLM endpoint and generate a professional report with PoCs and remediation steps.

Frequently Asked Questions about ai-threat-testing

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I test LLM applications for OWASP Top 10 vulnerabilities like prompt injection?

To test LLM applications for OWASP Top 10 vulnerabilities like prompt injection, you need a structured penetration testing workflow that systematically exploits risks, captures execution artifacts as evidence, and generates actionable remediation guidance.

What is the best way to assess RAG security and vector database poisoning risks?

Assessing RAG security and vector database poisoning risks requires running targeted offensive tests against your RAG pipeline to identify data poisoning vulnerabilities, capture proof-of-concept evidence, and document findings with risk scoring for hardening.

Can I use AI security testing to generate proof-of-concept evidence for model extraction attacks?

Yes, AI security testing generates exploit evidence by capturing prompts, outputs, and logs during authorized model extraction attacks, allowing you to produce reproducible proof-of-concept findings for your security assessment.

How do I run a focused LLM vulnerability assessment for a production chatbot before release?

To run a focused LLM vulnerability assessment for a production chatbot, apply specialized agent workflows to test specific OWASP LLM threats across full or focused scopes, generating prioritized vulnerabilities with PoCs and hardening guidance.

Does LLM penetration testing cover supply chain attacks and monitoring bypasses?

LLM penetration testing covers supply chain attacks and monitoring bypasses by executing specialized agent workflows designed to exploit these OWASP LLM Top 10 vulnerabilities and capture the resulting incident evidence.

What are the limitations of AI security testing for LLM endpoints?

AI security testing for LLM endpoints is limited to authorized penetration testing scopes, meaning it must only be applied to chat assistants, API endpoints, and RAG systems where explicit testing permission has been granted.