detoxify

Classifies text for toxicity, severe toxicity, obscenity, threats, insults, and identity-based attacks.

2|Updated Jan 15, 2026
One-click install
npx skills add https://github.com/DTMC-marketplace/governance --skill detoxify
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: detoxify
Source: https://github.com/DTMC-marketplace/governance/tree/main/skills/detoxify
Command: npx skills add https://github.com/DTMC-marketplace/governance --skill detoxify

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill helps identify and flag toxic comments and harmful content within text, contributing to safer online environments and compliance with content moderation policies.

Core Features & Use Cases

  • Toxicity Classification: Detects various forms of toxicity including general toxicity, severe toxicity, obscenity, threats, insults, and identity-based attacks.
  • Compliance Assessment: Assists in evaluating AI systems against regulatory requirements like Article 9 of the EU AI Act concerning societal risks.
  • Use Case: A social media platform can use this skill to automatically flag potentially harmful user comments for review before they are published, reducing the spread of hate speech.

Quick Start

Use the detoxify skill to analyze the provided text for toxicity.

Frequently Asked Questions about detoxify

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I detect toxic comments in user-generated content?

To detect toxic comments, you can use the detoxify model to classify text for various forms of toxicity including severe toxicity, obscenity, threats, insults, and identity-based attacks.

What types of harmful content can machine learning models classify?

Machine learning models for content moderation can classify harmful content including general toxicity, severe toxicity, obscenity, threats, insults, and identity-based attacks.

Can I use toxicity detection for EU AI Act compliance?

Yes, toxicity detection assists in evaluating AI systems against regulatory requirements like Article 9 of the EU AI Act concerning societal risks.

How do I analyze text for identity-based attacks and severe toxicity?

Analyze text for identity-based attacks and severe toxicity by passing the content through the detoxify classification model to receive scores across multiple toxicity categories.

Does content moderation require a specific model to detect threats and insults?

Accurate classification of threats, insults, and obscenity in content moderation requires the Detoxify model for precise detection of harmful user-generated content.

What is the best way to flag hate speech before it is published on a social media platform?

The best way to flag hate speech is to automatically analyze user comments with a toxicity detection model to identify harmful content for review before publication.