What problem does it solve? Verifying that a sandbox actually contains a hostile autonomous agent requires a live-fire test, and running, judging, and reporting that test involves many dispatch inputs, posture combinations, and verdict nuances that are easy to get wrong. ## Core Features & Use Cases - CTF Dispatch Guidance: Explains every workflow input for the breakout CTF suite, including model selection, monitor, sandbox backend (sbx or kata), auto mode, whitebox framing, and turn budgets. - Verdict Interpretation: Defines what CONTAINED, BREAKOUT, INCONCLUSIVE, and NO VERDICT each claim and, critically, what they do not claim, including posture-specific caveats. - Round Set Reporting: Describes how to read recorded results from the metrics-history branch and report them as a single table with run links, transcripts, and honest limitations. - Use Case: A maintainer wants to check whether the glovebox sandbox holds against a new model, so they dispatch five control-arm rounds, read the recorded verdicts, and report the containment results with transcript links. ## Quick Start Ask the assistant to run the breakout CTF with five rounds on a chosen model and summarize the containment results as a table.