aiml-nsfw-benchmark

Benchmark NSFW text classification across five categories using a standardized ISC template.

934|152|Updated Mar 1, 2026
One-click install
npx skills add https://github.com/wuyoscar/ISC-Bench --skill aiml-nsfw-benchmark
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: aiml-nsfw-benchmark
Source: https://github.com/wuyoscar/ISC-Bench/tree/main/templates/aiml_nsfw
Command: npx skills add https://github.com/wuyoscar/ISC-Bench --skill aiml-nsfw-benchmark

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

The NSFW benchmarking template provides a standardized, reproducible framework to evaluate DistilBERT NSFW text classification across multiple explicit content categories, enabling researchers to measure recall, coverage, and robustness in a controlled setting.

Core Features & Use Cases

  • Standardized NSFW categories (sexual_explicit, sexual_suggestive, harassment_sexual, grooming, sextortion) and language variants to test cross-language capabilities.
  • Ready-to-use prompts and sample datasets with deterministic evaluation flows, plus guidance for extending categories and anchors.
  • Use cases include model evaluation in AI safety research, benchmark-driven fine-tuning, and validation of prompt-variant experiments.

Quick Start

Run the included prompts and scripts to evaluate your NSFW classifier on the provided five-category benchmark.

Frequently Asked Questions about aiml-nsfw-benchmark

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark NSFW text classification across multiple categories?

You can benchmark NSFW text classification using a standardized ISC template to evaluate recall and category coverage across sexual_explicit, sexual_suggestive, harassment_sexual, grooming, and sextortion categories. The benchmark provides reproducible evaluation workflows with JSON samples and Python tooling.

What categories does the NSFW benchmark evaluate for AI safety research?

The NSFW benchmark evaluates five categories: sexual_explicit, sexual_suggestive, harassment_sexual, grooming, and sextortion. These standardized categories enable researchers to measure recall, coverage, and robustness in controlled AI safety evaluation settings.

Can I test cross-language NSFW detection capabilities with this benchmark?

Yes, the NSFW benchmark supports multi-language prompt sets and language variants to test cross-language NSFW detection capabilities. This allows evaluation of model performance across different languages for comprehensive AI safety validation.

How do I evaluate DistilBERT NSFW classifier performance with reproducible workflows?

You can evaluate DistilBERT NSFW classifier performance using the benchmark's deterministic evaluation flows with ready-to-use prompts and sample datasets. The included Python tooling and JSON samples ensure reproducible evaluation workflows for measuring recall and category coverage.

Does the NSFW benchmark support weak-anchor variant testing for prompt experiments?

Yes, the NSFW benchmark supports weak-anchor variants for prompt-variant experiments. This feature enables researchers to validate prompt-variant experiments and extend categories and anchors as needed for comprehensive model evaluation.

What's the best way to validate recall and category coverage in NSFW detection models?

The best way to validate recall and category coverage is using a standardized NSFW benchmark with reproducible evaluation workflows. The framework provides deterministic evaluation flows across five explicit content categories with JSON samples and Python tooling for measuring robustness.