Abstract
Evaluating LLM safeguards requires comparing refusal, attack success, and policy violation rates against real-world deployment risk rather than local scores alone.
Safeguards in deployed LLM services are evaluated by refusal, attack success, and policy violation rates. Those rates characterize how a control performed on the requests it was tested on. A deployment has to answer a different question: how much help with harmful tasks the service still gives an attacker who keeps adapting or finds another way in. We determine what each reported result implies for that question, allowing results from different safeguard families to be compared under one deployment criterion. The evidence requirements are strongly asymmetric. One attack that obtains harmful help from the deployed service suffices to establish that such help remains, and such attacks appear repeatedly in the coded record. Establishing that little remains cannot follow from the safeguard's own numbers alone; it also requires evidence about what the surrounding system still allows after the safeguard performs its local function. Such evidence is supported or derived in only a small minority of the depth-coded claims, and one such claim bounds its scoped residual. A better local score is therefore not, by itself, a stronger claim about the deployment. Safeguard research cannot stop at raising local scores; a gain has to be judged by whether it makes a deployed system any safer.
Community
This work fundamentally reexamines LLM safeguard research. It asks not only whether a safeguard works in evaluation, but whether it actually makes deployed systems safer, clarifying what evidence and research can translate into real-world safety gains.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs (2026)
- Decomposition Attacks Across Unlinkable Identities: Limits of Stateful Defenses for LLM Services (2026)
- Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models (2026)
- If Agents Were Angels, No Governance Would Be Necessary: Out-of-Band Policy Enforcement at a Trusted Tool Boundary (2026)
- Influence Is Not Authority: When Causal Guardrail Signals Make Legitimate Tool Use Look Like an Attack in Tool-Using LLM Agents (2026)
- SynthGuard-ReleaseBench: Locked-Audit Evidence for Synthetic Tabular Data Releases (2026)
- LMSM: LLM Security Framework Inspired by Linux Security Modules (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.00519 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper