SWARM Safety Eval

Compute SWARM soft-metrics from ledger data and diff against prior evals.

Updated Jun 3, 2026
One-click install
npx skills add https://github.com/swarm-ai-research/aeon --skill swarm-safety-eval
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: SWARM Safety Eval
Source: https://github.com/swarm-ai-research/aeon/tree/main/skills/swarm-safety-eval
Command: npx skills add https://github.com/swarm-ai-research/aeon --skill swarm-safety-eval

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

Grade the fleet's own multi-agent safety dynamics with SWARM soft-metrics (toxicity, quality gap, welfare), diff vs the prior eval, and notify on regression

Core Features & Use Cases

  • Compute SoftMetrics from ledger data
  • Diff against prior evals to detect regressions
  • Notify operators when degradation occurs

Quick Start

Run the evaluation on the current ledgers to generate a SWARM safety report and publish the article.

Frequently Asked Questions about SWARM Safety Eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I track multi-agent safety regressions across evaluations?

Multi-agent safety regressions are tracked by computing SWARM soft-metrics from ledger data and diffing results against prior evals. This flags degradation in toxicity, quality gap, or welfare, producing a machine-readable diff and a human-readable article.

What are SWARM soft-metrics and how do they evaluate fleet safety?

SWARM soft-metrics evaluate fleet safety by quantifying toxicity, quality gap, and welfare within multi-agent dynamics. They grade safety performance by analyzing ledger data to produce a comprehensive evaluation report and notify operators of any degradation.

How do I compute toxicity and welfare metrics from multi-agent ledgers?

Compute toxicity and welfare metrics from multi-agent ledgers by running a scheduled safety evaluation. The process integrates with the SWARM bridge and reporting workflow to extract SoftMetrics, diff them against prior evals, and flag regressions.

Does SWARM Safety Eval require the SWARM bridge and ledgers to work?

Yes, SWARM Safety Eval requires the SWARM bridge and ledgers to function. It integrates directly with these components to extract ledger data, compute soft-metrics, and generate reporting workflows for tracking multi-agent safety dynamics.

Can I automate multi-agent safety reporting on a schedule?

Automate multi-agent safety reporting by applying the evaluation on a schedule to compare current metrics against prior evals. This scheduled process flags regressions in toxicity, quality gap, or welfare, and automatically publishes a human-readable article.

What is the best way to notify operators about multi-agent welfare degradation?

Notify operators about multi-agent welfare degradation by applying scheduled safety evaluations that diff current soft-metrics against prior evals. This workflow integrates with reporting systems to alert operators when regressions in toxicity, quality gap, or welfare occur.