Source-linked AI summary

Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays

Bhaskar Gurram

arXiv:2608.14639v1cs.LGcs.AIcs.CL

TL;DR

Document-extraction systems need per-field accept/review decisions with controlled error among accepted fields, but natural procedures can violate that contract on real documents. This paper diagnoses those failures and develops a validity ladder, finding that rigorous document-level certification currently yields only 0.060 mean coverage on CORD while practical controls achieve 0.318 coverage at 0.096 risk.

  • Problem

    Per-field selective-risk procedures for document extraction lack validated evidence that accepted-field error stays below a target α on clustered real documents.

  • Method

    The paper diagnoses three failure modes and evaluates a validity ladder combining split protocols, Mondrian Learn-then-Test, exact binomial tails, and pre-specified provenance conditioning.

  • Results

    0.060 mean coverage is achieved by the rigorous document-iid PAC tier, while the practical tier attains 0.318 coverage at achieved risk 0.096.

  • Takeaways & Limitations

    Validity guarantees must be reported by tier: practical controls provide on-average risk control, whereas document-matched PAC certificates are honest but near-vacuous in this setting.

  • Takeaways & Limitations

    Useful rigorous coverage is demonstrated only on CORD×sonnet, while haiku and open-weights full-ladder replication remains queued and powered document-level PAC procedures remain open.

Abstract

from arXiv · show

Per-field accept/review with selective risk at most alpha -- accept a field only if the error rate among accepted fields is controlled -- is the trust contract document-extraction systems need, and the natural procedure silently violates it on real documents. On 13,859 genuine claude-sonnet-5 fields from 800 CORD receipts (49.0% correct) we diagnose three failure modes: document clustering (design effect 1.84-2.45), score-refit leakage (coverage 0.416 at risk 0.127, violating alpha=0.10 in 95% of splits), and a tie-mass pathology (a degenerate score collapses the threshold grid, 0.030 to 0.001). We organize the fixes as a validity ladder, guarantee form stated per tier. A fit/val split protocol restores expected-selective-risk control for a learned fusion: coverage 0.318 at risk 0.096 at nominal alpha=0.10, no tolerance band (production variant 0.326) -- an on-average point whose realized risk exceeds alpha in 47.5% of resplits, not a certificate. Mondrian Learn-then-Test with exact binomial tails yields per-group PAC certificates: field-iid 0.171 at risk 0.068, cluster-corrected 0.140, doc-iid 0.060 -- the only tier matching documents, honestly near-vacuous today. Support-bin, the pre-specified provenance taxonomy, wins every rigor tier on the sonnet CORD capture (p<1e-4, Bonferroni-corrected) -- a win that does not replicate on the same documents under haiku or qwen -- while on higher-accuracy corpora pooled thresholds win: conditioning helps exactly where pooled cannot certify, subsumed by a learned score elsewhere. A frozen-configuration confirmation on selection-untouched claude-haiku-4-5 held at both risk levels, and a blind three-annotator human-gold audit verifies the practical tier's accepted-set risk at 1.3% against its 10% budget (Fleiss' kappa=0.83; labels err one-sidedly pessimistic). Released Apache-2.0 with seed-pinned, regression-gated procedures.

1 Introduction

This paper shows that folklore per-field selective-risk procedures silently violate the trust contract on real document-extraction data, then provides diagnoses, protocols, certificates, and a characterization of when conditioning helps. Its experiments use genuine frontier-LLM outputs at scale and distinguish model-scoped empirical wins from corpus-general claims.

  • Problem: Selective risk is the error rate among accepted fields, which systems promise to keep below a target α.Each extracted field receives a confidence and accept/review decision under this per-field trust contract.
  • Contribution: The folklore procedure silently violates this contract on real documents, motivating three explicit levels of rigor.The paper presents what to run instead of simply fitting a confidence score and holding out a calibration split.
  • Experiment: 13,859 predictions from claude-sonnet-5 on 800 CORD receipts were only 49.0% correct, with additional FUNSD and XFUND-de captures spanning the difficulty spectrum.The testbed is deliberately real and difficult, with construction, labels, and measurement findings documented in a companion benchmark paper.
  • Diagnosis: Three failure modes are quantified: document clustering with design effect 1.84–2.45, score-refit leakage with risk 0.127 at nominal 0.10, and tie-mass pathology collapsing the threshold grid.The leakage violation occurred in 95% of splits, and each failure mode is pinned by a counterfactual experiment.
  • Characterization: Support-bin wins every rigor tier on the hard regime’s sonnet capture at p < 10−4, but the same win collapses under a weaker-signal model on identical documents.The paper characterizes a two-regime law: with a learned accept score, covariates belong in the score, while conditioning helps where pooled thresholds cannot certify.
  • Novelty: The paper claims no new conformal theory; its contributions are the diagnoses, protocol, certified application, and characterization.The underlying machinery is classical.

2 Related Work

Prior work supplies confidence scoring, selective-risk control, visual attribution, and evaluation foundations, but the paper adds grounding and risk-controlled accept/review guarantees for document extraction. Its procedures and guarantees are distinct from the companion benchmark’s datasets, labeling, and measurement studies, with reproducibility supported by a released harness.

  • Confidence for extraction: Existing extraction-confidence methods report calibration and selective-risk metrics or provide model-agnostic trust scores, but none supplies grounding or a risk-controlled accept/review guarantee.The cited methods include Beyond Logprobs, Cleanlab TLM, and real-time trustworthiness scoring.
  • Selective prediction and risk control: The paper builds on selective-classification bounds, conformal risk control, conformal factuality, and selective CRC for marginal risk control.Learn-then-Test reframes risk control as multiple testing; the paper’s rigorous tiers use exact binomial tails.
  • Companion papers: The paper owns procedures, guarantees, diagnoses, and characterization, whereas VerifyDocBench owns datasets, labeling, reliability auditing, and model/language measurement.Released materials support reproducibility through seed-pinned splits and regression-gated procedures.
  • Document-specific failure modes: Related work establishes visual attribution for risk-controlled generative OCR and VISA, while Traub and Kirchhof provide an evaluation perspective on selective prediction.These works motivate adjacent attribution and evaluation components for document-specific failure modes.

3 Setup

The setup defines field-level selective extraction as maximizing accepted coverage while keeping error among accepted fields at most α. Experiments use frozen text-layer Claude captures, five trust signals, alternative accept-score models, Mondrian conditioning, and document-level split evaluation.

  • Selective extraction contract: The trust contract maximizes coverage subject to selective risk—the error rate among accepted fields—being ≤α.Each field receives a confidence, grounding record, and accept/review decision; correctness is schema-typed, with omission and hallucination scored separately.
  • Data and captures: Headline experiments use genuine, frozen claude-sonnet-5 per-field outputs with k=3 self-consistency from OCR text, not page images.The haiku capture is reserved as selection-untouched confirmation data.
  • Signals and accept scores: Five signals combine self-report, self-consistency, grounding, entailment NLI, and ambiguity-penalized support, where support divides a match score by the number of equally good locations.The learned fusion adds engineered features and achieves test AUROC 0.925 vs 0.871 for the LR on CORD.
  • Conditioning: Mondrian conditioning applies the same threshold rule within pooled or taxonomy-defined groups, including the pre-specified support-bin provenance taxonomy.Candidate groupings also include fieldtype-freq and fieldtype-rule.
  • Evaluation protocol: 40 fixed document-level 50/50 calibration/test splits report mean selective risk, coverage standard deviation, and viol, the fraction of splits exceeding α.The protocol uses numpy generator seed 7, nominal α, no tolerance band, and sign-flip permutation tests with 20,000 flips for paired differences.

4 Why naive per-field selective guarantees fail on documents

The standard add-one selective-risk rule relies on exchangeability, but document-level extraction violates that premise through clustering, score-refit leakage, and discrete-score tie mass. These mechanisms can produce operating points that appear valid while substantially exceeding the nominal α=0.10 risk budget.

  • Combined failure: 0.114 coverage at 0.122 achieved risk, with 78% of splits violating nominal α=0.10, was the earlier draft’s “before” operating point.The operating point was produced by the diagnosed mechanisms plus taxonomy selection.
  • Failure 1: document clustering: 2.15 design effect on CORD, 1.84 on FUNSD, and 2.04 on XFUND-de reduce the effective calibration sample to roughly half its nominal size.At the add-one threshold, the design effect reached 2.45 at other thresholds; pooled CORD achieved risk was 0.105, violating in 50% of splits.
  • Failure 2: score-refit leakage: 0.416 coverage at 0.127 risk under same-half score fitting violated the nominal α=0.10 target in 95% of splits.The depth-3 gradient-boosted fusion fits its score and threshold on the same calibration fields, transferring score optimism into threshold selection.
  • Failure 3: tie-mass pathology: 1,702 distinct calibration values collapsed to 257 when an all-zero entailment column left only coarse discrete signals.Tie masses of 221 and 183 fields at the acceptance head force thresholds to accept a tie mass whole or not at all.

5 Procedures: a validity ladder

The validity ladder separates score–threshold independence from exchangeability: document-wise fit/validation splitting restores the former, while Mondrian Learn-then-Test adds per-group PAC certificates under explicit sampling assumptions. Its procedures control multiplicity and expose remaining limitations, including clustering, score-refit leakage, and conditioning-taxonomy selection.

  • Split protocol: Document-wise fit/validation splitting fits scores and transforms on fit data, thresholds and bin edges on untouched validation data, and never test data.This restores score–threshold independence but not exchangeability, so lower tiers remain marginal, on-average guarantees under document clustering.
  • Mondrian Learn-then-Test: Exact binomial-tail testing with Holm or fixed-sequence multiplicity control selects valid Mondrian thresholds, with the mix rule never collapsing across three corpora.Candidates are 15 geometric acceptance-fraction quantiles from 1%–100%, snapped to distinct-score boundaries; fine grids yielded pooled certified coverage 0.0009 vs 0.0055 on the 6.9k dump.
  • Guarantee ladder: With probability ≥1 −δ, tier 3 bounds each group’s true selective risk by α when within-group accepted-field errors are iid.The iid premise is load-bearing: design effect ≈2 motivates reporting ltt.neff, which deflates binomial n by the plug-in design effect; tier 4 replaces fields with documents.
  • Protocol limitations: Tiers 2–4 refit the 5-signal LR on the calibration data used for LTT p-values, formally reproducing the score-refit premise violation.A clean-fit control bounds the effect at ≈0 for this 5-parameter score, while the split-protocol twin reports 0.212 vs 0.218; camera-ready tiers 3–4 use the frozen fit-half score.
  • Conditioning selection: Pre-specifying support-bin avoids selecting the taxonomy by held-cell coverage, and every headline table reports that configuration rather than a per-cell winner.The support-bin-vs-pooled lift is evaluated with Bonferroni correction over four candidate taxonomies.

6 Main results

The main results establish a validity ladder in which practical and low-capacity methods control expected selective risk, while PAC tiers certify increasingly document-aligned guarantees at substantial coverage cost. Support-bin conditioning is strongest on sonnet CORD, whereas pooled thresholds win on easier corpora, and a frozen confirmation plus human-gold audit supports practical robustness.

  • Validity ladder: 0.318 coverage at 0.096 achieved risk is the practical tier’s headline, but realized risk exceeded 0.10 in 47.5% of resplits.The production no-NLI variant reaches 0.326 at 0.097; both are expected-risk operating points, not per-deployment certificates.
  • Validity ladder: 0.218 coverage at 0.095 is achieved by shared 5-signal logistic regression with add-one × support-bin under the split protocol.The threshold-half refit reaches 0.218 at 0.096 and is statistically indistinguishable from the split-protocol result.
  • Validity ladder: 0.171 coverage at 0.068 achieved risk is certified under field-iid PAC, while document-iid PAC provides only 0.060 mean coverage and certifies nothing in 19/40 splits.The field-iid assumption is violated here; document-iid certification bounds the macro per-document functional.
  • Conditioning: Support-bin wins every taxonomy comparison on sonnet CORD, including 0.171 versus 0.097 for fieldtype-rule at the LTT tier and 0.060 versus ≤0.006 for doc-LTT.All reported lifts are p < 10^-4 after Bonferroni correction; pooled LTT with the full Holm budget reaches 0.098 at 0.060 risk.
  • Generalization and audit: 0.491 coverage at 0.093 on FUNSD shows rigorous certification is attainable, but pooled LTT beats every taxonomy there; frozen haiku confirmation reaches 0.167 at 0.093.On FUNSD, the fieldtype-rule certificate reaches 0.128 at 0.034 with zero violations, while the haiku run held at nominal risk levels.
  • Generalization and audit: 1.3% accepted-set risk against blind human gold verifies the practical tier’s empirical performance under conservative automatic labels.Three annotators achieved Fleiss’ kappa=0.83; 21% of automatically flagged errors were actually correct and 0% were false-optimistic.

7 When does conditioning pay? A two-regime characterization

Conditioning pays in a two-regime pattern: learned scores should absorb covariates, while weak or frozen scores benefit from taxonomy conditioning when pooled thresholds cannot certify. Conditioning is harmful when the score already subsumes group information, and its benefit must exceed finite-sample penalties.

  • Regime 1: With a learned score, adding covariates inside the score beats external Mondrian conditioning, which either violates nominal risk or loses coverage.Field-type features add 0.069 coverage inside the score, while removing them for taxonomy conditioning recovers essentially none of that gain.
  • Regime 1: External conditioning on a full learned score fragments threshold samples because tree fusion already equalizes group score scales, imposing a per-group penalty without gain.The penalty scales like log(1/δ)/n_g for PAC tiers; fieldtype-Mondrian further reduces coverage by 0.062 for frequency and 0.036 for rule.
  • Regime 2: With a weak or frozen score, conditioning is the largest lever: shared-LR coverage rises from 0.134 pooled to 0.218 with support-bin, a +0.084 gain.On the 4-signal run, pooled 0.042 rises to fieldtype 0.133–0.135, approximately a 3× lift.
  • Regime 2: At the rigorous tier, conditioning rescues certification when pooled cannot, but fragmentation loses coverage when pooled already certifies: CORD support-bin 0.171 versus pooled 0.091–0.098.On FUNSD, pooled LTT reaches 0.280 versus 0.149 for support-bin; XFUND likewise favors pooled.
  • Where the NLI signal lives: The NLI signal is optional for learned tiers but essential for shared-fusion and rigorous tiers because it unties score values needed by the threshold grid.Backfilling produced 34.4% nonzero entailment and 2,016 distinct fused values; LR add-one coverage was 0.134 versus 0.037 with the dead column.

8 The price of validity

Real certification is substantially more conservative than the invalid add-one procedure on the 6.9k CORD dump. Under support-bin conditioning, retained coverage rises sharply with the allowed risk budget but vanishes at α=0.05 because the calibration sample is too small for a 90%-confidence certificate.

  • Coverage cost: 28% of invalid add-one coverage is retained at α=0.10 under the per-group PAC certificate.The comparison uses the 6.9k CORD dump and support-bin conditioning.
  • Coverage cost: 42% of invalid add-one coverage is retained at α=0.15, increasing the certified procedure’s practical viability.The certificate’s cost falls as the allowed risk budget grows.
  • Coverage cost: 84% of invalid add-one coverage is retained at α=0.20, but 0% is retained at α=0.05.Approximately 200 calibration documents cannot support a 90%-confidence certificate at α=0.05.

9 Limitations · A Guarantee statements and procedure details

The rigorous validity ladder provides different guarantees under exchangeability assumptions, but document clustering limits field-level validity, while nontrivial certification currently concentrates on CORD×sonnet and practical human-gold evidence remains limited.

  • 9 Limitations: 0.128 rigorously certified on FUNSD, but its taxonomy lift was not significant, while XFUND-de remains small-n.The full rigorous ladder certifies useful coverage on CORD×sonnet only.
  • 9 Limitations: The frozen-config haiku confirmation was the first out-of-selection replication; full haiku and open-weights ladders remain queued.This limits current evidence for generalization beyond the demonstrated corpus and extractor setting.
  • 9 Limitations: 1.3% human-verified accepted-set risk met the 10% budget, but scaling human gold to FUNSD and XFUND remains future work.Automatic labels were one-sidedly pessimistic.
  • A Guarantee statements and procedure details: Under calibration/test exchangeability, add-one marginal thresholds control E[selective risk] ≤ α over draws, but make no per-draw statement.The field is the exchangeability unit for Tiers 1–2.
  • A Guarantee statements and procedure details: Document-level clustering violates field exchangeability; the split protocol removes score–threshold dependence but does not resolve clustering.This distinguishes leakage control from the exchangeability assumption required by the marginal guarantee.
  • A Guarantee statements and procedure details: 0.103 risk occurred in 25% of splits for the grounded×support taxonomy under pure fixed-sequence testing, whereas the mix rule never collapsed across corpora.The failure mechanism was lucky-zero tiny-n bins with many groups.
  • A Guarantee statements and procedure details: Cluster correction replaces n_t with n_t/d_deff using estimated within-document error clustering, making the procedure approximate and uniformly more conservative.The design effect is estimated from per-document error clustering.

B Full disclosure: the entailment-capture defect and its forensics

The first 13,859-field CORD capture contained a silent NLI-stage failure that set entailment ≡0.0 for every field. This collapsed the fused score to 257 distinct values, producing anomalous certification behavior and severe tie masses.

  • Defect and root cause: 13,859 fields had entailment ≡0.0 because a mid-capture torch swap silently skipped the NLI stage.Per-field logs recorded entailment skipped alongside a torch-compile indexing error.
  • Forensic anomalies: ∼4× coverage shrinkage after doubling the data and a doc-level add-one of exactly 0 exposed the defect.
  • Score collapse: 257 distinct fused-score values remained when the only fine-grained continuous signal was dead, creating tie masses of 221/183 fields.

C Labeling reliability

Label stability is assessed by rescoring identical predictions under two automatic protocols, with a κ=0.10 caveat applying to every FUNSD free-text number.

  • C Labeling reliability: Table 9 rescored the same predictions under two automatic protocols and assigns every FUNSD free-text number a κ=0.10 caveat.

D Supplementary studies (simulated and floor-extractor)

Supplementary simulated and floor-extractor studies test provenance-conditioned conformal acceptance without placing these results beside genuine-model numbers. Conditioning recovers coverage at held risk when pooled acceptance collapses or cannot certify the grounded group.

  • Simulated controlled study: ≈0% pooled coverage versus 33–67% grounding-conditioned coverage, with risk held in every simulated condition at α=0.05.The simulated extractor used an uninformative accept score and 200 trials per condition.
  • Simulated controlled study: +0.50 mean coverage lift from provenance-conditioning at held risk in the simulated controlled study.Table 10 reports this lift at α=0.05.
  • Floor-extractor at-scale check: 0.24 →0.84 FUNSD coverage at achieved 2% risk and 0.01 →0.72 CORD coverage at held 10% risk using the floor extractor.CORD’s result required the ambiguity penalty; without the 1/m penalty, the grounded group was uncertifiable.
  • Floor-extractor at-scale check: The add-one guarantee holds tightly at large N, and conditioning lifts coverage where the two-regime characterization of §7 predicts.These are supporting checks using a conservative, low-recall label-search extractor rather than a frontier LLM.
Loading 2608.14639v1…