Source-linked AI summary

Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks

Volodymyr Ovcharov

arXiv:2608.15428v1cs.CLcs.AIcs.CY

TL;DR

Multiple-choice scores are only informative when models need the questions, but option-only solvability and answer-position habits can confound that measurement. This paper tests permutation-based, habit-corrected blind evaluation and finds that filtering items against one model does not neutralize them for a stronger model: GPT-5.6 still scores 0.515 on the survivors.

  • Problem

    Multiple-choice benchmarks lack evidence that models needed the question to identify the correct option, limiting score interpretation as domain competence.

  • Method

    The paper permutes option orders in a blind condition and separates each model’s answer-position habit from content-based accuracy, with selecting and reporting models held separate.

  • Results

    0.515: GPT-5.6 answered the question-hidden survivors correctly after Haiku 4.5 was filtered to 0.204 on them.

  • Takeaways & Limitations

    Habit-corrected blind baselines should accompany headline scores, and reference-style options avoid the problem observed with self-contained legal propositions.

  • Takeaways & Limitations

    The central demonstration uses one examination bank and one gating model, so it does not establish that every filter fails against every stronger model.

Abstract

from arXiv · show

Multiple-choice benchmarks are graded on whether a model picks the right option, not on whether it needed the question. Measuring that gap takes care: a model answering A to most items scores above chance wherever the key sits at A, and reads as recognition when it is not. We measure it on UA-JudgeExam: 11,990 four-option items with official keys, published by Ukraine's Higher Qualification Commission of Judges. Shown the options and no question, Claude Haiku 4.5 scores 0.383 against chance, and the leak is concentrated: 11.8% of items are answered blind on all eight option orders, against 0.2 items expected by chance. It is not quotation: search over 280,059 editions of Ukrainian legislation recovers 0.128. Gating those out retains 8,128 items, on which the gating model itself now scores 0.204, and GPT-5.6, which took no part in the selection, still answers 0.515 of them with the question hidden. Scoring twelve held-out models on the whole set and subtracting each one's answer-position habit, only two keep an excess: GPT-5.6 at +0.265, Sonnet 4.6 at +0.081. Without it the ranking misleads: Llama 3.1 8B scores 0.292 blind, above every model but those two, purely by answering A to 92% of items. The gate does select something real: on the items it rejected, eleven of twelve models score 0.518-0.789, every interval clear of what the same model scores on the items it kept. But that signal is one model's, and filtering on it does not transfer upward. Neither is visible on a 400-item sample, where nine models read as "statistically at chance". Rewriting distractors instead overshoots to 0.168, below chance and as exploitable. The same probe on LEXam returns chance: every option there points into the stem, none longer than 33 characters. Item format decides whether the problem can arise; capability decides how much is extracted. We release the corpus, the predictions and the harness.

1 Introduction · 2 Related Work

The paper treats question-hidden solvability as a validity threat in multiple-choice legal benchmarks and finds that filtering against one model does not produce a model-neutral set. It situates this problem alongside prior work on answer-without-question behavior, position bias, and annotation artifacts, while showing that option format—not legal subject matter—determines whether leakage arises.

  • 1 Introduction: A benchmark score evidences domain competence only when the question is load-bearing and the key cannot be identified from the options alone.Self-contained legal propositions can reveal which option states real law through authorities, deadlines, or statutory wording.
  • 1 Introduction: Filtering fails to transfer: the gating model scores 0.204 on survivors, while GPT-5.6 scores 0.515 on the same question-hidden set.The retained subset contains 67.8% of the original bank, and distractor rewriting instead falls below chance.
  • 1 Introduction: The proposed protocol separates content from answer-position habit by permuting options and forcing the key into each slot, making mean accuracy chance under content-independent choice.The selecting model is held separate from the reporting model; ten of twelve held-out models extract nothing under this estimator.
  • 1 Introduction: The probe returns chance on LEXam because its options are reference-style and no longer than 33 characters, indicating that option format determines whether leakage can arise.This contrasts with the self-contained propositions used in professional licensing material.
  • 1 Introduction: The paper releases UA-JudgeExam, a 11,990-item bank with official keys, independently verified extraction, the 8,128-item gated subset, and all predictions.The resource supports reproduction of the filtering and measurement analyses.
  • 2 Related Work: Prior work found above-majority choice-only performance across MCQA datasets and proposed metrics isolating the question’s contribution, motivating this paper’s blind solvability probe.Balepur et al. report success in 11 of 12 cases without evidence that memorization or individual-choice priors fully explain it.
  • 2 Related Work: Position-bias research shows that slot preferences and token bias can mimic above-chance accuracy, so the paper distinguishes answer-position effects from content using option-order permutations.A single fixed option order cannot tell whether a model recognizes content or favors the slot containing the key.
  • 2 Related Work: The work extends hypothesis-only artifact research to multiple choice and contrasts with legal judgment-prediction studies that examine outcome-revealing language in tribunal claims.Its focus is whether legal benchmark options themselves permit question-hidden recognition.

3 UA-JudgeExam

UA-JudgeExam is a four-option legal question bank published by Ukraine’s Higher Qualification Commission of Judges, with official keys for appellate-judge testing. Its self-contained legal propositions make answering from options alone a meaningful blind condition, unlike LEXam’s stem-dependent pointers.

  • Format: UA-JudgeExam’s options are self-contained legal propositions, whereas LEXam offers pointers such as “i and iii” into statements supplied by the stem.This difference makes the option-only blind condition meaningful for UA-JudgeExam but not for LEXam.
  • Dataset: 11,990 well-formed items were extracted from five Commission documents covering general knowledge and four legal specialisations, alongside 385 withdrawn and 40 invariant-failing items.The documents total 1,672 pages and come from Commission decision No. 221/зп-24 of 15 July 2024.
  • Extraction validation: 98.96% of question texts, 98.88% of option orders, and 98.41% of answer keys were confirmed across the full bank by an independent extraction path.A 200-item stratified pilot confirmed all three components in 0.995 of items; 192 full-bank items failed at least one check.
  • Extraction validation: 942 item-number cells were missing during extraction, but detecting a new stem amid accumulating options recovered the records without duplicate numbering within any source document.Naive extraction would merge adjacent items into one eight-option record.

4 The Solvability Profile

The benchmark contains little signal from surface heuristics or statutory quotation, but option-only performance reflects plausibility rather than verbatim recovery. Key-position imbalance is small yet measurable and is controlled explicitly.

  • Surface cues: No length or position heuristic clears 0.301, while the correct option averages 2.6 characters longer than distractors.The key appears at A in 28.9% of items, B and C in 24.8% each, and D in 21.5%.
  • Statutory quotation: 0.128 overall accuracy comes from searching 280,059 legislation editions, despite exact matches for the key in 0.580 of items and distractors in 0.438.When exactly one option matched, it was the key 65.1% of the time, showing that raw recoverability is a weak decision signal.
  • Statutory quotation: 94.6% of one- and two-token options occur verbatim, inflating raw match rates because bare figures match across many contexts.These options account for 10,637 of 47,960 options and can match for every item.
  • Statutory quotation: 0.433 versus 0.347, and 0.366 versus 0.414, are Haiku's and Sonnet's accuracies on verbatim keys versus other items in the 199-item pilot.The models disagree on the direction, every interval overlaps, and the paper concludes that the leak is plausibility rather than quotation.

5 The Blind Gate

The blind gate uses eight randomized option orders to identify items answered correctly without the question, revealing concentrated leakage rather than a uniform small effect. It retains 8,128 items, on which the gating model’s blind score falls to 0.204.

  • Gate design: The gate rejects items with at least five hits in eight trials, while requiring at least six parseable responses; it does not reject on the lower tail.Under the null, P(X ≥5) = 0.027, while zero hits occurs for 10.0% of clean items.
  • Gate design: 1,419 observed all-eight items versus 0.7 expected under seven effective trials indicates that repeated option orders slightly weaken the binomial reference but do not explain the concentration.The eight orders have 6.9 distinct orders on average, and all eight distinct orders occur only 27% of the time.
  • Leakage concentration: 1,419 items (11.8%) were answered blindly in all eight attempts, versus 0.2 expected under Binomial(8, 0.25), showing leakage concentrated in a minority.The pooled blind rate over the full bank was 0.383, and 3,517 items (29.3%) were never answered blind.
  • Selection outcome: 8,128 items (67.8%) were retained after the gate, from 8,137 items with four hits or fewer after nine failed the minimum-trials rule.Retention varied by specialisation: administrative 74.3%, commercial 70.6%, civil 64.6%, criminal 63.2%, and general 61.4%.
  • Selection outcome: 0.204 was the gating model’s blind score on the 8,128 retained items, making the bank clean by that model’s measure.Gating used Haiku 4.5, while reported blind figures came from models that took no part in selection.

6 Does the Repair Transfer?

Filtering a benchmark until one model reaches chance does not make it clean for stronger or different models: GPT-5.6 still answers 0.515 of Haiku-gated items blind, with a corrected excess of +0.265. The gate identifies real cross-model leakage, but its transfer is shaped by model-specific preferences and remains difficult to detect in small samples.

  • Transfer failure: 0.515 is GPT-5.6’s blind accuracy on items retained after gating against Haiku 4.5, whose own retained-set score is 0.204.Sonnet 4.6 scores 0.320 on the same items; corrected excesses are +0.265 for GPT-5.6 and +0.081 for Sonnet 4.6.
  • Position-bias correction: +0.265 and +0.081 are the only corrected excesses among held-out models; Llama 3.1 8B’s 0.292 blind score falls to +0.005 after accounting for its A preference.Llama answers A in 92% of items while A is the gold position in 29%, producing above-chance accuracy without corresponding leakage.
  • Cross-model signal: 0.518–0.789 is the blind-accuracy range for eleven of twelve held-out models on rejected items, and every model has a positive accepted-versus-rejected gap.The rejected items therefore capture leakage that generalizes across models, although Llama 3.1 8B scores 0.337 because of its near-universal A response.
  • Gate signature: 0.892 is the correlation between agreement with Haiku’s blind pick and the accepted-to-rejected gap, indicating that the gate carries the selecting model’s preferences.Agreement with Haiku ranges from 0.42–0.52 for ten of twelve models, versus 0.25 under independence.
  • Sample-size limitation: Eleven models appeared “at chance” in a 400-item sample, while nine extract nothing beyond positional habit once corrected.The small sample concealed both the limited genuine leakage and its dependence on the model that performed the selection.
  • Capability relationship: 0.916 is the Pearson correlation between full-condition and blind-accuracy rankings across eleven held-out models, but the association is carried mostly by two leaking models.For the nine remaining models alone, the association is r = 0.713 over a narrow blind-accuracy band.

7 A Benchmark That Does Not Leak, and Why

The blind probe returns chance on LEXam, where options point into the stem rather than state legal propositions. This contrast shows that option-only solvability depends on item format, not legal exams, jurisdiction, or language.

  • LEXam comparison: 0.250 chance: both models score at chance on LEXam, with the upper confidence bound below 0.250.GPT-5.6 recovers 0.515 of gate-accepted items from our bank blind but recovers nothing on LEXam.
  • Item format: 59.6% of our options exceed LEXam’s 33-character maximum, reflecting self-contained legal propositions rather than pointers into the stem.Figure 4 contrasts the banks’ option-length distributions and links the format difference to their blind-probe scores.
  • Item format: 1,655 LEXam items use options that reference statements enumerated in the stem, such as “i und iii” and “none of the statements”.The median option is 9 characters, and the longest is 33 characters; no legal proposition appears in an option position.
  • Implication: The result validates the probe and relocates option-only solvability from legal exams to item format.Self-contained legal propositions can be judged on their own, whereas pointer-style options cannot.

8 Remedies: One Clear Failure, Two Cautions

Alternative repairs do not produce a neutral item set: replacing distractors overshoots below chance, while model-written distractors create a possible self-recognition risk. Negation items are not shown to drive the effect, but the pilot is too small to settle that explanation.

  • One clear failure: None of the attempted repairs reaches chance, so predictable model error remains a property of the item set rather than a neutral benchmark.This conclusion applies to the alternative repairs tested in the section.
  • One clear failure: 0.168 blind accuracy replaces Sonnet 4.6’s 0.386 after real-law distractors, while full accuracy remains essentially preserved at 0.746 versus 0.792.The intervals for blind accuracy are disjoint (n = 197), so the repair overshoots rather than reaching chance.
  • One clear failure: No donor-selection rule lands on chance: similarity-based donors instead leave one-sided coherence or invert the artifact.Donors similar to the question make the key the odd one out; donors similar to the key invert the artifact.
  • Two cautions: 0.381 blind accuracy for Sonnet 4.6 on its own generated distractors exceeds 0.274 on human-written originals, but n = 113 leaves the result uncertain.The point estimate moves toward generator-family self-recognition, motivating the implementation caution against evaluating models that generated the distractors.
  • Two cautions: 0.250 blind accuracy on 36 negation items versus 0.416 on 161 others points away from negation driving the effect, but the intervals cannot separate them.The pilot is too small to rule the explanation either in or out.

9 Limitations

The study’s central demonstration is limited by its reliance on one benchmark and one gating model, incomplete full-scale re-evaluation, uneven model coverage, and missing human and broader benchmark baselines. These constraints limit how widely the findings can be generalized or interpreted.

  • Generalizability: One benchmark and one gating model do not establish that every filter fails against every stronger model; the authors recommend gating with several vendors.Filtering against Haiku 4.5 did not produce a clean set for GPT-5.6, although the procedure does not show that all filters fail universally.
  • Evaluation scope: Only blind scoring on accepted items was rerun across all 8,128 items; rejected-set and full-condition results remain based on a 600-item sample.The sample estimates have half-widths near ±0.05 and ±0.07, and model errors were correlated because all models used the same 400-item draw.
  • Coverage: 47.5% of Llama 3.1 8B full-condition calls produced no parseable letter, leaving its figure based on 202 of 400 items and excluding it from the correlation.DeepSeek R1 and Pixtral Large were budget-bound but recovered at 4,096 tokens; nine other models parsed above 99.5%.
  • Interpretation: No human blind baseline prevents determining whether Ukrainian lawyers would score near the gating model’s pooled 0.383, while the capability association rests on eleven models.The item-format claim is based on comparing two benchmarks, limiting what the design can establish about genre versus model effects.

10 Conclusion

The conclusion argues that blind benchmark scores require per-model answer-position correction and must be reported alongside question-visible scores. Filtering items using one model does not neutralize them for stronger models, while sample size and option design determine whether the problem is visible.

  • Measurement: 0.292 blind accuracy for Llama 3.1 8B on the gate-accepted set comes entirely from answering A on 92% of items.Its raw blind score exceeds every held-out model except the two that actually leak.
  • Filtering: 0.515 is GPT-5.6’s question-hidden accuracy on items retained after Haiku 4.5 was reduced to 0.204.GPT-5.6 took no part in selection, so filtering against one model does not neutralize the bank for a better one.
  • Evaluation scale: 400 items made nine models appear statistically at chance, while shared-sample errors moved together by 0.010 on the full set.The small-sample result concealed answer-position correction and correlated benchmark-wide movement.
  • Benchmark design: LEXam returns chance because its options point to statements in the stem rather than standing alone as self-contained legal propositions.The conclusion recommends publishing each model’s habit-corrected blind baseline beside headline scores when options are self-contained.

A Prompts

The benchmark prompts are Ukrainian, with options shown as A–D and permuted only in the gate; blind and full conditions differ by whether the legal question is hidden. Responses are parsed across several vendor formats, with unparsed outputs excluded from model-specific denominators.

  • Prompt format: Prompts are in Ukrainian, and options appear as A)–D) in the given order except during gate trials, where their order is permuted.The LEXam comparison uses the same two prompts in English because its split is German and English.
  • Prompt format: In the blind condition, the question is hidden and the model must select the most likely correct option using only the four answer choices.The required response is exactly one letter: A, B, C, or D.
  • Prompt format: In the full condition, the legal qualification-test question is shown and the model must choose the single correct option.The response format is likewise restricted to one letter: A, B, C, or D.
  • Response parsing: Unparsed responses are excluded from each model’s denominator rather than scored as incorrect.The parser checks explicit answer markers, a leading letter, and the last standalone letter; unparsed rates are below 0.5% for nine of thirteen models, with higher exceptions including Llama 3.1 8B at 25.5% and 47.5% in the full condition alone.

B Models

The study evaluates models accessed through Amazon Bedrock, including conventional systems spanning a wider capability range. Experimental settings were standardized where possible, with explicit exceptions for GPT-5.6’s temperature restriction and models’ differing token use.

  • Model access: All models were accessed through Amazon Bedrock in August 2026, with full snapshot identifiers reported because aliases may change.The release records complete snapshot identifiers for reproducibility.
  • Model coverage: Qwen3 32B, Gemma 3 12B, Ministral 8B, Nova Micro, and Llama 3.1 8B extend the conventional group toward lower capability.Three of these models overlap with Fan et al.’s small open-source group, making the sets partially comparable.
  • Experimental settings: GPT-5.6 omitted temperature because it rejected the parameter, while every other model ran at temperature 0.This preserves the provider-specific constraint rather than forcing an unsupported setting.
  • Experimental settings: The 600-item sweep used a 2,048-token output budget, whereas the full-scale blind run used 4,096 tokens.The budget affected models differently: Nova Pro was verbose, while DeepSeek R1 used tokens for reasoning traces absent from the response body.

Data and Code

The released materials cover the UA-JudgeExam corpus, gated subset, prediction sets, evaluation runs, controls, and extraction, verification, and gating code.

  • Released data: 11,990 corpus items and the 8,128-item gated subset are released with 600-item cross-vendor data and 16,800 blind and full predictions.The release also includes 9,600 first-sweep predictions and 7,200 from the small-model extension.
  • Released code and runs: 105,664-call blind runs, 21,600 prompt, labelling, and position-ablation calls, negative-result runs, reasoning controls, and supporting code are available at the stated Hugging Face dataset URL.The package includes extraction, verification, and gating code.
Loading 2608.15428v1…