What problem does it solve?
It helps you defend an artifact or plan against attacks from rival model outputs by requiring evidence-based rebuttals and concrete diffs instead of theater.
Core Features & Use Cases
- Live multi-model defense loop: Handles repeated “Round N” attack submissions with a consistent, structured output contract.
- Attack ledger with classifications: Normalizes attacks and labels them as Valid, Partial, Invalid, or Needs-More-Info to prevent aimless debate.
- Steelman-first rebuttals with plan_diff: Forces each response to state the strongest version of the attack, cite evidence, and propose concrete, artifact-scoped changes.
- Strict TACT evidence bar: Enforces Truth, Authenticity, and Clarity requirements and forbids hallucinated evidence or invented citations.
Quick Start
Run adversarial-defender to defend an artifact by pasting “Round 1” attacks and answering the five required Phase 1 inputs (artifact, rival model, evidence standard, environment access, exit condition) before the first attack.