Source-linked AI summary
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
Parsa Mazaheri, Kasra Mazaheri
TL;DR
Verifier wiring is treated as an engineering detail, but this paper tests whether prior audit–repair context changes what the checker reports. Across controlled ProcessBench evaluations, prior context consistently reduced false alarms without detectable discrimination gains, motivating caution about silent threshold shifts.
Problem
Whether audit–repair context changes a verifier’s reports is treated as an engineering detail despite its relevance to checking-pipeline design.
Method
The study compares byte-identical tasks after prior audit–repair episodes, measuring false-alarm rates on human-verified-correct traces against controlled contexts.
Results
Up to 11.5 pp fewer false alarms occurred across all 15 model × wording combinations, reflecting a criterion shift without detectable discrimination gain.
Takeaways & Limitations
Verifier wiring can silently alter thresholds, so reduced false alarms should not automatically be interpreted as improved checking.
Takeaways & Limitations
The evidence is bounded by model and wording coverage, nonseparated verdict and trace contrasts, and a single-author false-alarm audit without inter-annotator agreement.
Abstract
from arXiv · showhide
Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports. Measuring false alarms on human-verified-correct ProcessBench traces with the present task held byte-identical, we find that a completed audit -> repair episode already in the model's context lowers false alarms in 15 of 15 model x wording combinations, by 2.8 to 11.5 percentage points against a length-matched non-audit control, a 9 to 25% reduction relative to that control. The direction contradicts what the accumulated-message literature predicts: an episode whose audit reported an error lowers false alarms further still, at all five wordings on the model where that manipulation lands cleanly, though a negativity asymmetry predicts more flagging. Decomposing the episode finds repair content and audit verdict complementary: different components carry the effect on different model families. Signal-detection analysis locates the change in the threshold rather than in discrimination -- the criterion moves in 15 of 15 combinations and survives correction in 13 while d' survives in none, though the d' test is half as sensitive by construction -- and a hand audit of 50 false alarms finds 82% simply wrong, so at this operating point the shift need not be harmful. With reasoning enabled the effect keeps its relative size on both models tested, and the threshold reading holds there too.
1 Introduction
The paper argues that placing a verifier in a prior audit→repair context changes what it reports, contrary to the nearest literature’s prediction. The effect lowers false alarms without detectable discrimination gains, though removed false alarms are mostly wrong.
- Instrument: Up to 11.5 pp lower FAR occurs without a detectable discrimination gain, warning against treating flagged-item metrics alone as improvement.FAR measures errors reported in traces human annotators verified as correct, and each false alarm costs a repair cycle.
- Method: The design isolates context wiring from output format by holding the present task byte-identical and using a completed audit→repair exchange about another item.Rendered prompts were diff-tested across conditions.
- Effect: 15 of 15 model × wording combinations show lower false-alarm rates after a prior audit→repair episode than a length-matched filler, surviving five controls.The present task remains byte-identical across conditions, while the completed episode concerns a different item.
- Polarity drift: All five wordings on the cleanly manipulated model show that an episode reporting an error lowers false alarms further, opposite to the accumulated-message prediction.Repair content and the audit verdict are complementary across models, so no single component is necessary.
- Instrument: A hand audit shows the removed false alarms are mostly wrong, making the threshold shift beneficial at this operating point despite no detectable discrimination gain.The paper frames pipeline arrangement as an engineering choice that can alter verifier behavior.
2 Related work
Prior work shows that LLM judgments shift with conversational polarity and other contextual factors, while critique–correction research evaluates repair quality. This paper instead examines whether assigning the judge a subsequent job, including completed repair context, changes verification behavior.
- Judges and context: Temkit (2026) reports judgement drift toward preceding-conversation polarity across 84,088 calls to 12 models, with d = −0.17 and 1.52× more drift from negative histories.The effect concentrates on items where the model is uncertain at baseline.
- Verifier strictness: Zhou et al. (2026) establish verifier strictness as a movable ProcessBench axis, but do not test whether ordinary pipeline arrangements move it unintentionally.The present question concerns an effect arising without anyone intending to steer strictness directly.
- Judges and context: Prior judge studies document sensitivity to self-preference, sycophancy, example ordering, provided knowledge, and paraphrase.This work distinguishes its lever as the judge’s assigned next job rather than its interlocutor or wording.
- Critique–correction pipelines: CriticBench, CriticEval, and self-correction studies assess how well critique and correction are performed, whereas this work asks what completed repair context does upstream of those outcomes.Yang et al. (2025) decompose self-correction into confidence and critique components.
- Signal detection: Signal-detection analysis pairs false-alarm and labelled-incorrect-trace detection rates, reporting d′ and criterion c to distinguish discrimination changes from flagging reluctance.The passages identify this pairing as the minimum needed to interpret false-alarm-rate movement.
3 Method
The study tests audit-to-repair context effects using disjoint, source-stratified correct and incorrect ProcessBench arms, frozen self-generated episodes, and matched controls. Sampling, paired bootstrap intervals, and signal-detection measures are designed to separate false-alarm propensity from discrimination.
- Data and allocation: 929 clean targets, 50 warmup items, and 122 reserve items are allocated from 1,101 eligible all-correct ProcessBench traces in three disjoint, source-stratified arms.A reported error on the clean arm is therefore a false alarm by construction, and episode items cannot also be cold targets.
- Data and allocation: 929 labelled-incorrect traces form a second disjoint arm matched to the clean arm’s source mix, while 50 additional incorrect traces support AX and AXN episodes.Detection rate means returning the incorrect verdict, not localising the faulty step.
- Measurement and inference: FAR uses 8 samples per item at T=0.7, and both correct and incorrect arms feed d′ and criterion c so false-alarm shifts can be separated from discrimination changes.At T=0, hard 0/1 item outcomes could hide small changes in report propensity; reused episodes are handled with cluster bootstraps over 50 frozen episodes.
- Context manipulation: The manipulation inserts a completed, self-generated audit →repair episode before 465 of 929 clean targets, with frozen, byte-identical episodes and an internally re-measured no-context baseline.Each auditor generates its own episode at T=0, preserving the model-specific context while holding the target request identical.
- Context manipulation: AS, AO, US, and UO cross assistant-versus-user placement with self-versus-peer attribution, while length-matched labels and stripping tests enforce textual identity without placement confounding.The four-cell design addresses the interpretability problems of relabelling an assistant turn or moving content into a user turn alone.
4 A prior audit–repair episode lowers false alarms
A completed audit–repair episode lowers false alarms relative to length-matched filler context across all 15 model × wording combinations. The effect is specific to audit history, depends on both audit and repair components, and is not reproduced by an explicit leniency instruction.
- 4 A prior audit–repair episode lowers false alarms: −2.8 to −11.5 pp: false alarms fell in all 15 model × wording combinations against length-matched filler, with every p at the 20,000-replicate floor.At the primary wording, effects were −4.00, −3.59 and −8.83 pp across the three models; this result survived every applied control.
- 4 A prior audit–repair episode lowers false alarms: +1.58, −0.18 and −0.30 pp: AF −R0 was null on two of three models, showing the effect was specific to prior audit context rather than context generally.On Qwen3.6-27B, AF −R0 moved opposite to the episode, making that model’s episode −R0 figure conservative rather than inflated.
- 4 A prior audit–repair episode lowers false alarms: −1.31, −2.35 and −3.13 pp: at F1, AS −AN survived on all three models, whereas AN −AF survived on one; across five wordings, survival reversed to 12 of 15 versus 10 of 15.The component ordering therefore depends on wording: neither audit nor repair is dispensable, and neither is safe to interpret from one wording alone.
- 4 A prior audit–repair episode lowers false alarms: −3.59 pp versus −1.76 pp: on Qwen3.6-35B-A3B, the episode moved FAR more than a maximal explicit leniency prime, whose [−3.72, +0.07] interval touched zero.This comparison cuts against a simple instruction-following account: the prior audit effect was established, while the explicit instruction was not.
5 Polarity drift alone cannot explain it
The error-verdict manipulation lowers rather than raises false alarms on Ministral, contradicting a polarity-drift account, but the test is clean on only one of three models. Component analyses likewise find model-dependent contributions without identifying one common mechanism.
- Polarity drift test: −5.62 pp [−8.27, −2.87] at the primary wording, with negative effects at all five wordings, shows error-verdict episodes lower Ministral false alarms.The five-wording range is −4.05 to −5.90 pp, and every result survives its declared family.
- Polarity drift test: 37 of 50 Qwen3.6-27B repairs end with a “verdict”: “correct” assertion, making its contrast null at every wording; Qwen3.6-35B-A3B survives only at F1.The 35B-A3B F1 effect is −3.83 pp [−5.51, −2.23], so the wedge is counted as one model of three.
- Interpretation: AX −AS rules out polarity drift and verdict echoing, but not broader semantic priming or base-rate calibration because AX also varies content and difficulty.A content-matched verdict flip would require a different design, since forcing the verdict would violate the invariant that episodes are the model’s own work.
- Component decomposition: Each component fails to survive on the model the others explain, so the evidence supports different model-specific routes rather than one common mechanism.All identified components move in the same direction across models, but no single route is established.
6 A threshold move, and a beneficial one
The audit-repair episode shifts models toward flagging less, without a detectable discrimination gain. Because most baseline false alarms are wrong, this leniency improves the operating point here.
- Threshold versus discrimination: The criterion moves toward flagging less in 15 of 15 model × wording combinations and survives correction in 13, while ∆d′ survives in 0 of 15.The median ratio of |∆c| to |∆d′| is 1.85, with sensitivity estimates positive in all but two combinations.
- Operating point: Balanced accuracy rises in 15 of 15 combinations and survives correction in 6, including +1.22 pp [+0.38, +2.14] on the 27B and +2.62 pp [+1.14, +4.10] on Ministral at F1.Detection falls by 0.3, 1.6 and 1.2 pp, but false-alarm reductions are 2.8, 3.5 and 6.5 pp.
- False-alarm validity: In a hand audit of 50 R0 false alarms from Ministral at F1, 41 (82%) are simply wrong, 8 (16%) defensible-but-stricter, and 1 (2%) a suspected goldlabel error.The 16% defensible share has a Wilson interval of [8.3%, 28.5%].
- False-alarm validity: Defensible flags concentrate in logical errors at 7/20, with none in 21 arithmetic or algebraic cases; arithmetic and algebraic flags account for 97% of the FAR reduction.The pooled comparison survives with Fisher exact p = 0.0034, while neither error type alone survives correction over ten pairwise comparisons.
- Conclusion: The episode makes models more lenient without a detectable discrimination gain, and four in five baseline false alarms being fabrications makes leniency beneficial at this operating point.A lower false-alarm rate alone cannot distinguish reluctance to flag from improved discrimination, so detection rates on 929 incorrect traces provide the second arm.
7 Robustness
Robustness checks show that the audit-repair effect is not safe to interpret at one wording and persists with reasoning enabled, where it remains a threshold shift rather than improved discrimination. Reproducibility controls and prospective-obligation results further delimit the finding’s scope.
- Wording: All eight tested contrasts require wording sweeps because the pre-specified wording was repeatedly unrepresentative, making single-wording interpretations unsafe.Table 2 covers four claim contrasts, while Table 4 covers all eight.
- Reasoning traces: −1.30 pp and −1.10 pp: the filler-controlled AS −AF effect survives reasoning on both Qwen models.Relative reductions are −19.7% and −17.5% with thinking on versus −21.6% and −14.3% with it off; the absolute margin shrinks as the baseline improves.
- Reasoning traces: +0.087 and +0.058: ∆c moves away from flagging, while −0.0046 and +0.0069: ∆d′ remain on zero with reasoning enabled.Balanced accuracy does not rise in this setting, with all four p ≥0.16.
- Reproducibility: 0.625960: AS reproduced its false-alarm rate exactly in a three-month re-run, while R0 reproduced to five decimals on the same frozen pools.Content hashes prevent contrasts from crossing pools because T=0 is not reproducible under continuous batching.
- Prospective responsibility: 13 of 13 surviving contrasts: every prospective repair-obligation rung that moves d′ on Qwen models moves it down, unlike the experienced episode.Because the prospective ladder splits by model family, it is reported in Appendix B rather than treated as part of the main robustness claim.
8 Conclusion
Pipeline wiring changes checker reports: a completed audit →repair episode lowers false alarms across all measured model × wording combinations, in a direction opposite to prior-context polarity drift predictions. Because the shift is an unrequested threshold change, its benefit depends on the operating point silently altered by the pipeline.
- 8 Conclusion: Up to 11.5 pp: completed audit →repair context lowered false alarms in every measured model × wording combination.The effect therefore appears across the full set of tested combinations.
- 8 Conclusion: Opposite to prior-context polarity drift predictions: where the manipulation landed cleanly, an episode reporting an error lowered false alarms further.The direction is not one practitioners should rely on without checking the resulting operating behavior.
- 8 Conclusion: An unrequested threshold move can help or harm depending on the operating point silently altered by the pipeline.The conclusion cautions that the effect’s value cannot be assessed independently of that operating point.
Limitations
The evidence is limited by model coverage, unresolved component contrasts, and outcome-associated dropout in the reasoning arm. The wedge is not consistently separable across models or AX conditions.
- Model and wording coverage: The wedge holds at all five wordings on only one of three models, while the 35B reaches its family at one wording.Qwen3.6-27B’s AX cell is degenerate rather than contradictory, with 37/50 repairs still asserting correct.
- Model and wording coverage: 37/50 repairs in Qwen3.6-27B’s AX cell still assert correct, making that cell degenerate rather than contradictory.
- Component identification: AXN and AN draw on different pools, so neither contrast separates the verdict token from the trace it describes.
- Reasoning-arm coverage: The reasoning arm covers only two of three models, and its truncation and schema-invalid-output dropout is outcome-associated rather than random.
A The instrument-sensitivity screen
The screen retained models that responded to framing relative to model-specific sham noise bands, rather than specifically to leniency. Qwen3.6-35B-A3B passed, while Gemma candidates did not, and leniency effects varied across retained models.
- The instrument-sensitivity screen: Three framing controls—PC, PCL, and PCH—were judged against each model’s own sham-derived noise band before hypotheses were tested.A model passed if any control interval cleared its band in either direction.
- The instrument-sensitivity screen: Qwen3.6-35B-A3B passed upward-only because PC shifted +6.78 pp and PCH +12.28 pp, whereas PCL yielded −1.76 pp [−3.72, +0.07].The screen measured responsiveness to framing overall, not leniency specifically.
- The instrument-sensitivity screen: 1.94 pp was Gemma’s framing range versus 8.54 pp for the 35B, and Gemma candidates cleared none of the three controls and were screened out.Their generated cells remained visible for selection auditing, but no claim rested on the screened-out pair.
- The instrument-sensitivity screen: −13.93 pp [−16.89, −11.14] on Ministral, −4.07 pp [−6.16, −2.28] on Qwen3.6-27B, and −1.76 pp [−3.72, +0.07] on Qwen3.6-35B-A3B were leniency-prime FAR shifts.Because the downward prime did not resolve on the 35B, a downward null there is weak evidence; measured effects are reported rather than a detectability bound.
B The prospective ladder … F The false-alarm audit rubric
Across the prospective ladder and matched audit-context comparisons, attribution is null while conditionality shows a wording-stable effect, and Ministral’s prior verdict changes false-alarm rates within the same trace pool. The instrument uses byte-identical audit tasks with controlled prior-context variants, while the false-alarm rubric distinguishes wrong, defensible, unclear, and label-error cases.
- B The prospective ladder: +2.78 to +3.23 pp is the wording-stable conditionality effect on Qwen3.6-27B across five wordings, with the sign opposite to work-avoidance predictions.The passage identifies conditionality as the most wording-stable effect in the study.
- D A within-pool verdict split: −10.30 pp [−13.48, −7.04] is the Ministral AS −R0 contrast when the prior episode reported an error, versus +0.77 pp [−2.75, +4.23] when it did not.On Qwen3.6-27B, the corresponding split is flat at −2.91 versus −2.74 pp; the comparison is observational.
- C Every claim contrast at every wording: Table 4 lists all eight claim contrasts at every wording across all three models, reporting false-alarm-rate percentage points and survival counts.Negative values mean fewer false alarms; F1–F5 are the five semantically matched audit wordings, with F1 pre-specified.
- E The instrument: AF is a per-episode length-matched filler control whose request and reply name neither correctness, error, review, nor repair.The replacement restates a held-out problem at a word count computed from the audit episode it replaces, with targets from 37 to 637 words.
- E The instrument: R0 uses only the target request, while the other three conditions prepend a two-turn held-out exchange before the byte-identical target request.AS contains the model’s audit and repair, AN replaces repair with an audit restatement, and AV deletes the second turn.
- F The false-alarm audit rubric: The false-alarm rubric classifies cases as wrong, defensible, unclear, or label_error, using an unweighted proportional sample from 390 flagged items.One author performed the audit, with no inter-annotator agreement; per-case verdicts and reasons were released for re-checking.
G Reproducibility details
The study used locally served, open-weight models under fixed inference settings and reproducibility seeds. Frozen episode pools were content-hashed, contrasts checked for matching content, and analyses recorded item-level outputs and intervals.
- Models and serving: Two additional Gemma candidates were screened out, while Qwen3.6-27B, Qwen3.6-35B-A3B, and Ministral-3-14B-Instruct-2512 comprised the served revisions.The passage identifies each model by its served revision hash.
- Integrity checks: Every contrast verified identical content hashes for its paired frozen episode pools before formation.Multiplicity families and screen membership were declared in code, with warnings for incomplete family populations.
- Integrity checks: The analysis recorded per-item flags, per-item confidences, and each contrast’s item-only interval alongside its clustered results.Incomplete declared families were tested at α/k using the smaller populated k, producing a more lenient threshold than declared.