safety-evaluation

Evaluate AI model safety across harm categories using rubrics and thresholds.

6|Updated May 30, 2026
One-click install
npx skills add https://github.com/jassics/awesome-claude-security --skill safety-evaluation
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: safety-evaluation
Source: https://github.com/jassics/awesome-claude-security/tree/main/plugins/ai-safety/skills/safety-evaluation
Command: npx skills add https://github.com/jassics/awesome-claude-security --skill safety-evaluation

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill addresses the need for comprehensive safety evaluation of AI models, providing a structured framework to measure and enhance safety across various harm categories.

Core Features & Use Cases

  • Structured Evaluation: Conducts safety evaluations structured across harm categories such as disallowed content, robustness, and truthfulness.
  • Rubrics and Thresholds: Implements clear rubrics and pass/fail thresholds for evaluating AI system performance.
  • Reporting: Generates detailed safety evaluation reports to inform release decisions and track regression over time.

Quick Start

Run the safety-evaluation skill with the provided prompts and evaluate the AI system's response against established harm categories.

Frequently Asked Questions about safety-evaluation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate AI model safety across harm categories?

To evaluate AI model safety, you can use a structured framework that assesses under-refusal, over-refusal, robustness, groundedness, and bias. It applies predefined harm categories with clear rubrics and pass/fail thresholds to measure system performance.

What is included in an AI safety assessment for robustness and truthfulness?

An AI safety assessment for robustness and truthfulness includes evaluating the model against disallowed content, measuring over-refusal and under-refusal rates, and testing groundedness. It uses predefined rubrics to generate detailed safety evaluation reports.

How do I set pass/fail thresholds for AI safety evaluation?

You can set pass/fail thresholds for AI safety evaluation by establishing clear rubrics for each harm category, such as robustness and bias. These thresholds determine if the AI system passes the safety assessment and generates reports for release decisions.

Does AI safety evaluation require specific tools for measuring over-refusal and bias?

Yes, comprehensive AI safety evaluation requires specific expertise and tools to accurately measure under-refusal, over-refusal, and bias. The process uses predefined harm categories and rubrics to ensure accurate safety assessment and regression tracking.

What's the best way to track AI safety regression over time?

The best way to track AI safety regression is by conducting repeated safety evaluations using consistent harm categories, rubrics, and pass/fail thresholds. This generates detailed reports that inform release decisions and highlight performance changes over time.

When do I need a structured AI safety evaluation framework?

You need a structured AI safety evaluation framework when preparing an AI model for release and wanting to measure safety across harm categories like disallowed content, robustness, and truthfulness. It provides the rubrics needed to track regression over time.