model-extraction

Simulate AI model extraction attacks and log defense responses.

3|Updated Nov 18, 2025
One-click install
npx skills add https://github.com/pluginagentmarketplace/custom-plugin-ai-red-teaming --skill model-extraction
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: model-extraction
Source: https://github.com/pluginagentmarketplace/custom-plugin-ai-red-teaming/tree/main/skills/model-extraction
Command: npx skills add https://github.com/pluginagentmarketplace/custom-plugin-ai-red-teaming --skill model-extraction

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill helps security teams evaluate how resilient AI models are to extraction attempts by simulating controlled probing and logging defenses.

Core Features & Use Cases

  • Structured attack techniques: Provides query-based extraction, distillation, embedding theft, and architecture probing scenarios for risk assessment.
  • Detection-focused evaluation: Includes indicators and metrics to measure fidelity, surrogate behavior, and defense effectiveness.
  • Real-world use case: Run authorized security tests against deployed models to identify vulnerability patterns and strengthen safeguards.

Quick Start

Run the extraction test script to simulate attacks and generate a report.

Frequently Asked Questions about model-extraction

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I test AI model resistance to extraction attacks?

To test AI model resistance to extraction attacks, you can simulate controlled probing scenarios including query-based extraction, distillation, embedding theft, and architecture probing to evaluate defense responses.

What metrics are used to measure surrogate model fidelity during security testing?

Surrogate model fidelity during security testing is measured using detection rates, embedding similarity, and surrogate behavior metrics to evaluate defense effectiveness and vulnerability patterns.

Can I evaluate detection mechanisms for embedding theft and distillation attacks?

Yes, you can evaluate detection mechanisms for embedding theft and distillation attacks by logging defense responses and analyzing structured detection indicators during simulated extraction attempts.

What's the best way to assess vulnerability patterns in deployed AI models?

The best way to assess vulnerability patterns in deployed AI models is running authorized security tests that apply configurable attack techniques and generate structured reporting metrics.

Does this model extraction testing support architecture probing scenarios?

Yes, model extraction testing supports architecture probing scenarios alongside query-based extraction, distillation, and embedding theft to provide comprehensive risk assessment for deployed AI models.

When do I need to run extraction attack simulations for risk assessment?

You need to run extraction attack simulations for risk assessment when evaluating how resilient deployed AI models are to unauthorized extraction attempts and strengthening safeguards against identified vulnerabilities.