aiml-fraud-benchmark

Evaluate fraud-detection pipelines against a scam-script corpus and category taxonomy.

934|152|Updated Mar 1, 2026
One-click install
npx skills add https://github.com/wuyoscar/ISC-Bench --skill aiml-fraud-benchmark
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: aiml-fraud-benchmark
Source: https://github.com/wuyoscar/ISC-Bench/tree/main/templates/aiml_fraud
Command: npx skills add https://github.com/wuyoscar/ISC-Bench --skill aiml-fraud-benchmark

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

ISC template for evaluating AI safety workflows by benchmarking fraud-script recall and validation against a classifier.

Core Features & Use Cases

  • Provides a standardized scaffold with categories, anchor data, and evaluation scripts to measure fraud recall.
  • Validates category coverage, absence of placeholders, and presence of manipulation tactics in scripts.
  • Use Case: teams can run automated checks to compare model outputs against expected fraud patterns across categories.

Quick Start

Run the benchmark using the provided scripts and classifier interface to start evaluating fraud recall.

Frequently Asked Questions about aiml-fraud-benchmark

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark fraud detection recall for LLM safety classifiers?

Benchmark fraud detection recall by automating evaluation against a predefined scam-script corpus and category taxonomy. This skill provides standardized scaffold scripts to measure recall, validating category coverage and manipulation tactics across multiple fraud scenarios.

What does an automated fraud text classification benchmark evaluate?

An automated fraud text classification benchmark evaluates category coverage, absence of placeholders, and presence of manipulation tactics in scripts. It compares model outputs against expected fraud patterns to validate safety pipelines and ensure consistent results.

Do I need a custom dataset to run scam detection evaluation scripts?

You need an accessible dataset of scripts and a classifier interface to run scam detection evaluation scripts. The skill provides anchor data and a category taxonomy, but reproducible reporting requires connecting your accessible dataset to the automated evaluation pipeline.

Can I use this benchmark to validate manipulation tactics in scam scripts?

Yes, you can validate manipulation tactics in scam scripts using the automated evaluation checks. The benchmark validates the presence of manipulation tactics and category coverage to ensure your fraud detection pipeline correctly identifies expected scam patterns.

How to generate reproducible reports for fraud recall validation?

Generate reproducible reports for fraud recall validation by running the provided evaluation scripts against a classifier interface. The standardized scaffold ensures consistent results by validating category coverage and comparing model outputs across multiple fraud categories.

What is the best way to compare model outputs against expected fraud patterns?

The best way to compare model outputs against expected fraud patterns is using automated benchmark scripts with a category taxonomy. Teams run checks to validate category coverage and absence of placeholders, producing consistent evaluation results across fraud scenarios.