A New Severity Scale Moves Agent Red-Teaming Past Binary Attack Success
Source: arXiv — 2026-07-08
Summary
Harry Owiredu-Ashley argues that binary attack-success-rate metrics used in agentic red-teaming discard exactly the information defenders need — how bad a successful attack actually was. The paper proposes a 7-level (L0–L6) ordinal harm rubric scored on reversibility, scope-crossing, and privilege escalation, evaluated via both a deterministic oracle and a frontier-LLM judge panel.
Key Takeaways
- Binary "did the attack succeed" metrics can't distinguish a trivial, reversible slip from an irreversible, scope-crossing compromise.
- The L0–L6 severity scale grades harm by reversibility, whether the action crossed a permission/scope boundary, and privilege escalation.
- Combines a deterministic scoring oracle with an LLM-judge panel for cases the oracle can't automatically classify.
- Gives defenders a concrete, comparable severity metric for tool-using agent red-team results, not just a pass/fail count.
Discussion
(No questions yet — ask follow-ups via a Claude Code chat session on this repo; answers get appended here.)