Source-linked AI summary

The Safeguard Worked. Is the LLM System Safer?

Pingyu Wu, Weiming Zhang, Nenghai Yu

arXiv:2609.00519v1cs.CRcs.AI

TL;DR

The paper asks how much harmful assistance remains after safeguards, beyond what local refusal or attack metrics show. It converts reported results using their outcome and attacker class into deployment conclusions, finding that evidence for residual assistance is asymmetric and that better local scores alone do not establish greater deployment safety. Its scope is limited by fixed-suite rates and tested-modification samples that cannot bound all adaptive histories or reachable attacks.

  • Problem

    Local safeguard metrics characterize tested requests, but deployment assessment must determine how much harmful assistance remains under adapting attackers or alternate paths.

  • Method

    The paper maps each reported result to deployment conclusions using its measured outcome and attacker class, with an auditable coding instrument recording required evidence.

  • Results

    The evidence requirements are asymmetric: 108 of 152 wide-coded claims report residual harmful assistance above zero, while only 5 of 24 depth-coded claims supply the least common fact needed to establish how little remains.

  • Takeaways & Limitations

    A safeguard can work as tested while deployed-system safety remains unresolved; evaluation gains should therefore be judged by the deployment conclusion they support.

  • Takeaways & Limitations

    Fixed-suite error rates do not constrain every adaptive history, and tested modifications are inner samples that provide no upper bound over all reachable attacks without exhausting the strategy space.

Abstract

from arXiv · show

Safeguards in deployed LLM services are evaluated by refusal, attack success, and policy violation rates. Those rates characterize how a control performed on the requests it was tested on. A deployment has to answer a different question: how much help with harmful tasks the service still gives an attacker who keeps adapting or finds another way in. We determine what each reported result implies for that question, allowing results from different safeguard families to be compared under one deployment criterion. The evidence requirements are strongly asymmetric. One attack that obtains harmful help from the deployed service suffices to establish that such help remains, and such attacks appear repeatedly in the coded record. Establishing that little remains cannot follow from the safeguard's own numbers alone; it also requires evidence about what the surrounding system still allows after the safeguard performs its local function. Such evidence is supported or derived in only a small minority of the depth-coded claims, and one such claim bounds its scoped residual. A better local score is therefore not, by itself, a stronger claim about the deployment. Safeguard research cannot stop at raising local scores; a gain has to be judged by whether it makes a deployed system any safer.

1 Introduction

The paper argues that local safeguard results do not directly answer how much harmful assistance remains in deployment. It develops an evidence-based conversion from reported evaluations to the strongest deployment conclusion they support.

  • Evidence gap: Adaptive attacks can exceed 90% success against most defenses despite original evaluations reporting rates near zero.Jain et al. measured 0 to 1% first-turn success rising to 5.4 to 14.0% after 15 rounds of adaptation.
  • Evidence gap: Deployment assessment must determine how much harmful assistance a guarded service still supplies, not merely compare safeguard techniques or benchmark scores.Existing surveys, guidance, safety-case work, and uplift analyses address related steps but do not replace this deployment question.
  • Method: The paper reads each result against its measured outcome and attacker class, then derives the strongest deployment conclusion that the evidence supports.If a reported quantity supports no nontrivial deployment conclusion, reanalysis cannot create one; a different measurement or missing system property is required.
  • Results: 108 of 152 wide-coded claims report an adverse value putting residual harmful assistance above zero.By contrast, establishing how little remains requires three deployment facts together, and only 5 of 24 depth-coded claims supply the least common fact: what remains possible after the check succeeds.
  • Takeaway: The coded record treats benchmark gains as hypotheses about deployment safety rather than guarantees.The supported conclusion follows the evidence reported, not the safeguard technique category.

2 Definitions and Scope

The paper defines deployment safety around residual harmful assistance from a safeguarded service under a fixed evaluation anchor. It distinguishes retained assistance from safeguard effect and requires rights and utility constraints alongside any residual bound.

  • Definitions and scope: Deployment safety concerns the harmful assistance a focal service still supplies after safeguards act, subject to basic-rights and legitimate-use constraints.It is a property of an intervention in a specified deployment, not an intrinsic property of an isolated mechanism.
  • Residual assistance: Z measures harmful assistance supplied through outputs, actions, and state transitions, normalized to [0,1], rather than harm ultimately achieved by an attacker.Lower Z values are less adverse.
  • Evaluation anchor: The evaluation anchor fixes the service, intervention, attacker-strategy class, rights constraints, utility requirement, harmful-assistance functional, and operating conditions.Its scope covers only the attacker strategies and operating conditions for which the conclusion is asserted.
  • Residual assistance: The primary deployment quantity is residual assistance Z(S[D]), while safeguard effect VZ(D) alone leaves that residual undetermined.The two quantities separate removed assistance from retained assistance.
  • Decision criterion: A zero-residual certificate corresponds to τ = 0, but deployment decisions must also verify rights constraints and the benign utility requirement q.Refusing service or excluding legitimate users cannot satisfy the criterion by itself.
  • Decision criterion: Because Z is normalized, every deployment begins with Z(S[D]) ∈[0,1]; evidence matters by supplying a tighter bound.The next section derives which published evidence supports each bound.

3 What a Reported Quantity Can Bound

The paper converts reported safeguard measurements into the tightest sharp bounds they support on residual harmful assistance, Z(S[D]), under fixed outcome and attacker coordinates. It shows that deployment-wide coverage, reachability, continuation, and dependence assumptions—not local scores alone—determine whether a nontrivial or zero-residual conclusion follows.

  • The Deployed Quantity: Z(S[D]) measures the harmful assistance a guarded service still supplies across every strategy available to the declared attacker.A local score describes one test, whereas the deployment quantity ranges over the attacker class.
  • Evidence Asymmetry: One observed attack establishes residual harmful assistance, but showing that little remains requires evidence covering the full attacker class.This asymmetry follows from the supremum defining the deployment quantity: lower bounds need one attainable law, while upper bounds must cover every law.
  • Noninformative Evidence: A fixed-suite marginal error rate supports no simulation bound because it averages over histories rather than constraining the conditional kernel after each history.This row is therefore noninformative regardless of how low the fixed-suite rate is.
  • Closed Mediation: Even with ε = 0, the sharp bound can equal one when α = 0 or r = 1, so flawless local checks can remain compatible with maximal deployment risk.The conversion quantities α and r belong to the surrounding deployment rather than the classifier.
  • Closed Mediation: A zero-residual mediation certificate requires complete coverage, zero conditional failure, and no continuation after covered success.These are deployment-wide facts that a local score does not report.
  • Composition: Composition improves bounds only under history-uniform conditional failure limits; with merely marginal bounds, perfectly correlated failures make any stack no better than its strongest layer.Alternative paths and retries can still accumulate attempts at a fixed per-attempt rate.

4 Reading the Literature Against the Schedule

The paper translates reported safeguard evidence into sharp, auditable deployment bounds by coding the anchor, required slots, and evidential status of each claim. The resulting schedule separates lower and upper residual bounds, verdicts, and unresolved cases while exposing exactly which evidence is missing.

  • A claim instance requires a deployed intervention, an operationalized adverse outcome, and either a comparison world or a formal connection between them.
  • Each coordinate is sourced, derived by a declared rule, or left unknown, and unknown coordinates receive no default.
  • The ten coding slots divide into four attacker-retention slots and six deployment-control slots, with endpoints produced only when every consumed slot is supported or validly derived.
  • The coding instrument makes incompleteness diagnostic because an open endpoint identifies the exact missing fact required by its construction.
  • The lower slot block establishes floors on residual harmful assistance, while the upper block establishes ceilings; unsupported constructions leave the corresponding endpoint open.
  • Only supported or validly derived relations enter endpoint computation, while claimed, unreported, or out-of-scope relations do not.
  • The resulting interval establishes a met tolerance when Ux ≤τx, an exceeded tolerance when Lx > τx, and otherwise leaves the deployment unresolved; claim verdicts remain separate.

5 What Published Safeguard Evidence Establishes

Published safeguard results establish deployment conclusions only within the attacker, outcome, and interaction conditions they measure. Positive residuals are repeatedly witnessed by attacks, whereas zero-residual certificates require jointly supported coverage, failure, and continuation evidence that is rarely available.

  • 5.1 One Coded Set at Two Coding Depths: Only five of 24 depth-coded instances support or validly derive evidence about what remains reachable after a locally successful safeguard event.This continuation evidence is the least frequently supplied of the three facts needed for an upper bound through mediation.
  • 5.1 One Coded Set at Two Coding Depths: Only one of those five instances supports all three gates under one anchor, making it the subset’s sole computable upper endpoint through mediation.Supporting continuation alone, or any other operand in isolation, does not tighten the endpoint.
  • 5.2 The Attack-Witness Row, and How the Literature Reaches It: Attack-witness values directly provide lower bounds on residual harmful assistance, but only for the strategies actually run within the declared attacker class.This bound does not require the safeguard’s local failure score.
  • 5.4 Restriction-Only Patterns: Restriction-only evaluations show improvement at selected operating points without establishing that the best attainable adverse value has moved at matched normal utility.None of seven capability-removal instances establishes a frontier change under that condition.
  • 5.3 Coverage and Continuation Determine a Check’s Deployment Bound: A 93% defense-success rate does not close a deployment bound when earlier harmful outputs, fresh-session retries, and permitted follow-on actions remain outside the measured event.The unreported continuation term leaves the upper endpoint open at one.
  • 5.6 Where the Three Gates Close: Fides supports a scoped zero-residual certificate by jointly establishing complete coverage, zero conditional failure, and zero continuation, with task-completion loss up to 24.5%.Structural separation supplies the continuation condition, while the utility cost remains a deployment trade-off rather than part of the certificate itself.

6 Relation to Prior Systematizations

Prior systematizations make safeguard mechanisms, attacks, and evaluation resources comparable, but comparable reported quantities do not by themselves establish what residual harmful assistance remains. This paper adds an auditable conversion from source-anchored measurements to the strongest deployment conclusion those measurements support.

  • 6 Relation to Prior Systematizations: Prior surveys, reviews, taxonomies, and SoKs differ in which step from safeguard mechanism to deployment decision they hold fixed.Each step contributes a different part of the path, including shared vocabulary, threat-model declarations, and broader models under explicit assumptions.
  • 6 Relation to Prior Systematizations: Comparable reported quantities with explicit provenance still leave open how much harmful assistance a guarded service supplies.An identically computed attack-success rate does not automatically become a residual claim, even under a declared threat model.
  • 6 Relation to Prior Systematizations: The paper converts each reported quantity into the strongest residual conclusion supported on its declared outcome scale and attacker class.The conversion also records which coding slots prevent a stronger conclusion, making comparisons auditable across safeguard families.

7 Discussion

Deployment claims depend on the quantities and premises required by their evidentiary route, not on isolated improvements in local safeguard scores. Different residual-risk gates therefore require different interventions and explicit scope, continuation, and dependence assumptions.

  • Evaluation Requirements: An evaluation should fix the anchor, identify the schedule row supporting its conclusion, and report every quantity that row requires.Improving one reported number cannot tighten a bound when another required quantity is missing.
  • Evaluation Requirements: Coverage is an architectural fact established by enumerating deployment paths through the mediation domain, including paths that reset scoped state.Accuracy figures do not establish coverage.
  • Evaluation Requirements: The residual continuation value r must include released content, permitted actions, accumulated state, and further attempts within the horizon.Because Equation 24 scales accuracy gains by 1−r, omitting r leaves the value of an improvement undetermined.
  • Evaluation Requirements: Layered error rates may be multiplied only when each conditional bound remains valid after every preceding interaction history at the deployed history grain.The dependence premise and its grain matter more than the number of layers.
  • Transferable Artifacts: For transferable artifacts, an upper-bound route requires a declared tampering class containing Σ and a uniform value bound over that class.Testing additional individual attacks adds lower-bound witnesses, not the required uniform bound.
  • Design Implications: Lowering ε, raising α, and driving r to zero are respectively statistical, architectural, and structural interventions.Detection accuracy alone cannot drive r to zero.
  • Design Implications: For external-world outcomes where the dual-use floor binds, τ is a deployment target constrained by L and the utility requirement B(g) ≥q.Choosing τ is a deployment decision, not an evaluation result.
  • Scope: A certificate transfers to another deployment only if its anchor and row premises are preserved or a containment argument covers the new scope.The certificate is not an intrinsic property of an isolated mechanism.

8 Conclusion

Deployment safety concerns the harmful assistance a guarded service still supplies, rather than whether a local safeguard worked on a specified test. Establishing little residual assistance requires broader evidence about coverage, conditional failure, and post-success continuation.

  • 8 Conclusion: A single successful attack establishes that harmful assistance remains, whereas showing little remains requires coverage, conditional failure, and continuation evidence together.One coded claim supplies this conjunction and rules out the worst case within its scope.

Ethical Considerations

The paper evaluates how published safeguard results support deployment-risk conclusions for the stakeholders who produce, assess, and act on those reports. It presents an evaluation standard rather than an attack capability.

  • Ethical Considerations: The study recomputes published quantities from public literature without executing attacks, using deployed systems, or involving human subjects.Its full-text coding was performed by two independent model channels on public papers.
  • Ethical Considerations: The stakeholders are safeguard-evaluation authors, reviewers and evaluators, and operators who act on evaluation reports.The paper's schedule identifies which reported quantities bound deployment risk and which provide no bound.

Open Science

The supplementary materials document the review tasks, source-channel records, and per-source assessments behind the paper's coded results. The appendix supplies formal assumptions and proofs establishing the stated bounds and their sharpness.

  • Supplementary Materials: The first supplementary file records the common full-text review task applied by both independent model channels.The remaining files document source-channel eligibility, claim-instance status, instance counts, and final per-source assessments.
  • Supplementary Materials: The supplementary materials are available in the project repository.Third-party full texts are not redistributed, but assessment files retain source identifiers, locators, and substantiating excerpts.
  • Formal Basis: The appendix assumes probability laws on a declared measurable trajectory space, with standard Borel spaces and finite or countable structure only where stated.It defines PT as a pushforward under the declared value-relevant projection and [x]+ = max{x,0}.
  • Formal Basis: For measurable f ranging in [0,1], the appendix uses elementary expectation facts and the fact that projection cannot increase total variation.These facts support the subsequent propositions.
  • Formal Basis: Proposition 1 follows by evaluating the supremum at a law in Cg, while Proposition 7 follows from inclusion of one feasible-law class in another.Equality holds when the two classes coincide.
  • Formal Basis: Testing finitely many modifications cannot upper-bound WZ(g) below 1 unless the tested set is shown to cover the relevant strategy class.An untested strategy can preserve observed coordinates while placing all mass on a trace with vZ = 1.
  • Formal Basis: The realizable frontier is ordered against the outer frontier because the latter minimizes over a superset, while restricting the feasible set cannot lower the optimum.The argument applies without identifying the two law sets.
  • Formal Basis: Pointwise positivity is insufficient on an infinite space to establish the uniform dual-use floor.With vZ(tn) = 1/n and b(tn) = 1, every trace has positive adverse value but the frontier at q = 1 has infimum zero.

A.3 Value-Relevant Simulation

The section derives deployment bounds from value-relevant simulation and shows that marginal average error rates cannot establish such bounds. Tightness results identify when local controls can support nontrivial or zero residual conclusions.

  • Value-Relevant Simulation: The total-variation coefficient in the value bound is sharp, so no smaller uniform coefficient follows from the stated premises.A two-point construction makes both the expectation difference and total variation equal to δ.
  • Value-Relevant Simulation: A simulation certificate requires conditional error bounds after every shared history, not merely a marginal error rate over a fixed suite.The sequential coupling argument yields a product bound when each adaptive evidence update remains close under every relevant history.
  • Value-Relevant Simulation: A fixed-suite average rate, however low, supports no simulation bound because an attacker can steer into histories where the conditional distance is large.The constructed example has arbitrarily small benign average disagreement but total-variation distance one on the attacker-selected histories.
  • Value-Relevant Simulation: Two deployments with the same benign trajectory law and maliciously reachable-law set agree on the relevant value, worst-case, and simulation functionals.The result follows because each functional depends only on those two objects.
  • Value-Relevant Simulation: A zero-residual certificate through mediation requires complete path coverage, zero conditional failure, and no continuation after covered success.These three conditions are jointly necessary and sufficient at the strict boundary.

A.6 Robust Trusted State

The robust trusted-state analysis bounds residual risk using the nearest maliciously reachable state distribution and conditional simulation. It shows that averages, labels, and layer counts do not substitute for deployment-relevant reachability or history-uniform evidence.

  • A.6 Robust Trusted State: The robust trusted-state bound depends on d⋆, the distance to the nearest maliciously reachable state marginal, rather than average acquisition accuracy.An allowed path can preserve d⋆ and the bound even while the average success rate over many paths tends to zero.
  • Composition Bounds: History-uniform conditional bounds license multiplicative composition without requiring independence, whereas marginal bounds reduce any stack to its strongest layer.Perfectly correlated failures attain the marginal minimum bound.
  • Composition Bounds: Alternative paths accumulate attempts under the union bound, while retry bounds follow from multiplying per-attempt conditional factors.The resulting bounds are attained by disjoint alternative-path events and sequential Bernoulli trials.
  • A.6 Robust Trusted State: For additive outcomes, coordinate-wise frontier bounds sum when the marginal optimizers admit an admissible joint coupling.The conclusion does not extend to union or intersection payoffs, and subclass bounds combine only when the subclasses cover the full attacker class.

B.3 Screening, Eligibility, and Machine Assistance

The review pipeline combines conservative screening, authority and eligibility criteria, independent machine-assisted coding, and randomized full-text sampling. Its records preserve anchors, evidence states, derivations, endpoints, and separate claim verdicts.

  • Screening Pipeline: The executed pipeline orders keyword identification, an authority gate, dual-channel screening, human quality checks, and randomized full-text sampling.Version-family deduplication reconciles records without adding an eligibility criterion.
  • Screening Pipeline: Title-and-abstract records advance only when both independent channels include them; disagreement or uncertainty excludes them from the full-text pool.A detected quality problem triggers rerunning the preceding screening step.
  • Eligibility: A claim instance requires a locatable intervention and adverse outcome, plus at least one additional anchoring condition; missing coordinates remain unknown.Sources without a codable instance remain in corpus denominators rather than being dropped.
  • Coding Strata: Truncation changed the positive-residual share from 71.1% to 70.8%, within the estimate’s 4.5-point sampling error.The corresponding counts were 108/152, 100/141, 89/124, and 80/113.
  • Coding Records: The coding record stores fixed anchors, ten evidence slots, derivations, source locators, endpoints, residual conclusions, and an independent claim verdict.Residual conclusions and claim verdicts are explicitly allowed to diverge.
  • Coding Strata: The wide-coded stratum pooled 152 claim instances from 187 distinct papers, while the depth-coded subset contained 24 instances.The two channels sampled the full-text pool independently, so wide-coded instances were pooled across channels rather than deduplicated by paper.

C.5 Case Record I: All Three Gates Closed

The Fides case is the depth-coded subset’s sole zero-residual instance because its integrity-scoped outcome closes all three gates under the declared attacker class. The record also preserves substantial performance and scope boundaries.

  • Case Record I: The Fides instance is the only depth-coded case with both endpoints computable and the only one receiving a zero-residual conclusion.Its anchored claim concerns untrusted data influencing consequential tool actions.
  • Case Record I: The coded assessment derives rx = 0 from the binary integrity outcome and supports the required coverage, failure, and continuation premises.The route-specific integrity value is fully identified even though some general anchor coordinates remain unknown.
  • Case Record I: The resulting normalized bound is 0 ≤ Zx(Sx[Dx]) ≤ 0, yielding a zero-residual certificate and an upheld claim verdict.The certificate applies within the declared attacker class and matching consequential-action semantics.
  • Case Record I: The design incurs up to 24.5% policy-on task-completion loss and two to three times the Basic planner’s token use plus query_llm latency.These costs do not enter the endpoint but must be weighed against the deployment’s declared target.
  • Case Record I: The certificate covers integrity-scoped outcomes with trusted configuration, not text manipulation, implicit confidentiality leakage, or all prompt-injection attacks.Those broader outcomes belong to sibling instances.
Loading 2609.00519v1…