Hermes Wiki
AIDigest/2026/07/12/2026-07-12-06-action-graded-severity-scale-agent-security

A New Severity Scale Moves Agent Red-Teaming Past Binary Attack Success

Source: arXiv — 2026-07-08

Summary

Harry Owiredu-Ashley argues that binary attack-success-rate metrics used in agentic red-teaming discard exactly the information defenders need — how bad a successful attack actually was. The paper proposes a 7-level (L0–L6) ordinal harm rubric scored on reversibility, scope-crossing, and privilege escalation, evaluated via both a deterministic oracle and a frontier-LLM judge panel.

Key Takeaways

  • Binary "did the attack succeed" metrics can't distinguish a trivial, reversible slip from an irreversible, scope-crossing compromise.
  • The L0–L6 severity scale grades harm by reversibility, whether the action crossed a permission/scope boundary, and privilege escalation.
  • Combines a deterministic scoring oracle with an LLM-judge panel for cases the oracle can't automatically classify.
  • Gives defenders a concrete, comparable severity metric for tool-using agent red-team results, not just a pass/fail count.

Discussion

(No questions yet — ask follow-ups via a Claude Code chat session on this repo; answers get appended here.)

Hermes Wiki