media-voicefake-benchmark

Generate SSML-based voice deepfake attack samples across five attack categories.

934|152|Updated Mar 1, 2026
One-click install
npx skills add https://github.com/wuyoscar/ISC-Bench --skill media-voicefake-benchmark
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: media-voicefake-benchmark
Source: https://github.com/wuyoscar/ISC-Bench/tree/main/templates/media_voicefake
Command: npx skills add https://github.com/wuyoscar/ISC-Bench --skill media-voicefake-benchmark

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This template provides a structured benchmark for evaluating voice deepfake detection pipelines using SSML-formatted samples across five attack categories, ensuring realistic, controlled test data.

Core Features & Use Cases

  • SSML-based attack samples to assess robustness of detection systems.
  • Voice-attack coverage (voice cloning, emotion manipulation, identity spoofing, ransom deepfake, and authority impersonation) for comprehensive testing.
  • Reusable framework for security research and benchmark development, with an explicit replacement step for PLACEHOLDER_SPEECH to simulate credible social-engineering scripts.

Quick Start

Replace PLACEHOLDER_SPEECH with realistic SSML content across all attack categories and run the validation workflow.

Frequently Asked Questions about media-voicefake-benchmark

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark voice deepfake detection pipelines using SSML?

You can benchmark voice deepfake detection by generating SSML-based attack samples with realistic content. This framework evaluates detection robustness across five social-engineering categories, ensuring all generated samples meet minimum length and category-specific marker requirements.

What types of voice deepfake attacks can I test for detection robustness?

You can test detection robustness across five attack categories: voice cloning, emotion manipulation, identity spoofing, ransom deepfake, and authority impersonation. Each category uses SSML-formatted samples containing specific social-engineering markers to evaluate comprehensive detection coverage.

How do I generate realistic social-engineering scripts for deepfake voice testing?

To generate realistic scripts, replace the PLACEHOLDER_SPEECH variable with credible social-engineering content across all attack categories. This explicit replacement step ensures the SSML blocks simulate authentic voice attack scenarios for security research and benchmark development.

Does this voice deepfake benchmark require specific detection APIs or dependencies to run?

No specific detection APIs or dependencies are required to generate the SSML attack samples. The framework provides a reusable, standalone template for security research, allowing you to apply the generated samples to any external voice deepfake detection pipeline you are evaluating.

Why do my generated voice deepfake SSML samples fail validation?

SSML samples fail validation if they do not meet the minimum length requirement, lack category-specific social-engineering markers, or contain invalid SSML blocks. Ensure all PLACEHOLDER_SPEECH variables are properly replaced with realistic content to pass the validation workflow.

What is the best way to test audio deepfake detection coverage for authority impersonation?

The best way to test authority impersonation detection is to use SSML-based attack samples containing realistic social-engineering scripts. This benchmarking approach specifically evaluates whether your detection pipeline can identify authority impersonation markers within the generated voice content.