Source-linked AI summary
Generative Gap Filling
Yonathan A. Arbel, David A. Hoffman
TL;DR
Contract scholarship has assumed that text provides little guidance once an express term is missing, leaving courts to rely on defaults, context, or policy. The paper tests that assumption by masking negotiated terms in real contracts and asking humans and language models to reconstruct them. Humans recovered terms substantially above chance, while models were correct 88% of the time, supporting text-based evidence as a supplement to gap filling while leaving AI-authored contracts outside the study’s foundation.
Problem
Contract scholarship has assumed that the remaining text provides weak evidence about an unstated bargain, making gap-filling methods difficult to validate.
Method
The authors mask negotiated terms in real executed contracts and compare predictions from lay readers, legally trained readers, and language models against the known originals.
Results
88%: language models recovered masked contract terms, outperforming human groups; approximately two-thirds of success reflected general deal-structure expectations and one-third specific language.
Takeaways & Limitations
Text-derived model predictions can serve as ordinary, contestable evidence for interpreting incomplete contracts and may reduce reliance on free-ranging normative gap filling.
Takeaways & Limitations
The paper’s proxy for contractual intent assumes contracts are written by people; AI-agent-drafted contracts may break that foundation.
Abstract
from arXiv · showhide
Most contract litigation turns on contracts that imperfectly record parties' bargains. When the parties' dispute can't be solved by interpreting the text, courts fill the gap. Scholars have long assumed that the remaining text runs out quickly, and provides thin evidence of the actual deal on the disputed point. On that view, a judge who supplies the missing term must be drawing on something else, from commercial defaults to her own policy preferences. Despite generations of work, courts have no real alternative to such unruly methods. We tested that assumption. Taking real contracts, we masked a term the parties had negotiated and asked readers to predict what we removed. Lay respondents recovered the hidden term about half the time, twice what chance predicts. Law students and lawyers did marginally better. But large language models, given nothing but the rest of the contract, recovered it nearly nine times in ten. The deal, in short, testifies to far more of the agreement than the literature assumes, including terms the parties never wrote. A contract, we argue, is like a radio signal from far away. Even when incomplete, enough of the message is carried elsewhere that the missing part can be reconstructed with the right receiver. True gaps are rarer than supposed. Courts can weigh model predictions as ordinary, contestable evidence, and parties can discipline the practice with "Choice of Model" clauses.
INTRODUCTION
Contracts often contain more information about imperfectly expressed bargains than gap-filling theory assumes. The authors test this by masking negotiated terms and comparing human and language-model predictions against known ground truth.
- Motivation: Contracts use inherited templates and elaborate drafting, yet can still appear unfinished when disputes arise.The introduction frames contracts as imperfect expressions of bargains despite efforts to make them robust to errors.
- Motivation: Courts supply missing terms when agreements do not directly address contingencies, distinguishing this work from ordinary interpretation.Examples include vanished indexes, destroyed music halls, and floating prices requiring judicial resolution.
- Method: The authors mask negotiated contingency terms in real contracts and ask laypeople, legally trained readers, and language models to reconstruct them.Because the original terms were negotiated, priced, drafted, and signed, the experiment supplies a known correctness benchmark.
- Results: Lay readers recovered hidden clauses about half the time, lawyers 59%, and language models 88% of the time.The models remained ahead of every human group, and differently trained or prompted models often returned the same answer.
- Results: Approximately two-thirds of model success came from general deal-structure expectations and one-third from specific contract-language inferences.Models were nearly perfect for commercially normal terms but right about 60% of the time for idiosyncratic outcomes.
- Implications: The findings challenge the premise that courts must rely mainly on defaults, context, or policy once express language ends.The authors argue that text-based AI evidence can support interpretation and help reduce reliance on normative gap-filling.
- Implications: A choice-of-model clause could fix the system used to interpret a contract and prevent post-dispute model selection from controlling the result.The proposal adapts the familiar contractual practice of selecting governing law and forum.
I. GAP FILLING’S EMPIRICAL GAP
Gap-filling scholarship differs over judicial methods but largely assumes that contract text eventually stops reliably evidencing the parties’ deal. The authors argue that this empirical boundary has been assumed rather than measured, and that true gaps are narrower than supposed.
- Shared premise: The literature treats the remaining text and context as second-best inferential sources when no express term addresses the dispute.This assumption supports longstanding concern that judges may substitute their own preferences for the parties’ intentions.
- Existing approaches: Gap-filling scholarship includes default rules, commercial context, formalism, policy-based fillers, and refusals to fill gaps.These approaches differ over what courts should do and what counts as a gap.
- Shared premise: Despite their disagreements, these approaches share the premise that the document eventually stops reliably evidencing the parties’ actual deal.Beyond that point, decisionmakers are thought to rely on negotiations, prior dealings, trade customs, penalties, or policy.
- Research question: The authors ask empirically where contract evidence actually gives out rather than resolving competing taxonomies of interpretation and construction.They begin from the limited assumption that a contract is partial evidence of a bargain, like any other witness.
- Research gap: The authors note that existing accounts do not establish how reliably courts can derive terms from text without express language.The literature invokes normative theories to justify intervention while leaving the practical deduction method underdeveloped.
- Research question: They distinguish true gaps, where silence reflects unresolved disagreement, from false gaps, where the surrounding document still supports recovery of a term.Their claim is that false gaps occupy less territory than the literature supposes.
FILLING
The paper tests whether contracts contain enough contextual evidence to reconstruct negotiated terms that are not visible in the remaining text. Humans perform above chance, while language models recover masked terms substantially more accurately, though performance varies by scenario and the claim has defined limits.
- Human respondents: 55% of lay respondents correctly recovered hidden contract terms, compared with chance prediction at half that rate.Performance varied from 32.2% in the artist scenario to 71.1% in the contingency fee scenario.
- Human respondents: Law students performed marginally better than lay respondents, while lawyers predicted missing terms correctly nearly 60% of the time.Lawyers outperformed other human groups on Artist and Contingency Fee but not Bottles.
- Language models: 88.3% of LLM unmasking predictions were correct, substantially exceeding human performance across the scenarios.Models surpassed 90% on Artist and answered every Contingency Fee question correctly, while Bottles was more challenging.
- Language models: Frontier-model accuracy ranged from 70.0% for GLM 5.1 to 100% for Opus 4.6, with several models between 88% and 97%.Qwen 3.6 reached 82%, while Gemini 3.1 Pro, Grok 4.2, and GPT 5.4 ranged between 88% and 97%.
- Robustness and interpretation: On challenging disputes, supplying the contract increased winner-prediction accuracy from 26% to 49%, but this exploratory analysis has limited evidentiary scope.The hard-dispute classification depended on the other two models’ answers, and the preregistered advance difficulty screen showed no reliable overall gain.
III. THE PRACTICE AND PERILS OF GENERATIVE GAP FILLING
Generative gap filling is presented as a useful but contestable aid to adjudication, not a replacement for judges. The authors propose Choice of Model clauses while emphasizing reliability, legitimacy, and scope concerns.
- Judicial Use: Generative gap filling should supply evidence for judges to weigh through ordinary adversarial procedures, reasoning, and appeal—not replace judicial decisionmaking.The authors reject oracle-like deployment and preserve open, contestable judicial evaluation.
- Choice of Model Clauses: Choice of Model clauses would specify the model, harness, and prompt protocol governing disputes about contractual silence.The clause fixes the model before disputes arise, addressing concerns about post-dispute model selection.
- Choice of Model Clauses: Model panels may improve contestability and replication, but correlated errors prevent them from providing independent verification.A strict-majority panel returned no answer on eight of 119 contracts, requiring aggregation and abstention rules.
- Drafting Effects: Pre-testing selected models before signing could reduce inadvertent contractual gaps and make remaining silences more likely deliberate.The authors envision both parties jointly observing model predictions about contractual silences before execution.
- Judicial Role: Choice of Model clauses do not eliminate normative judicial work; they provide a different evidentiary base for deciding whether model output reflects the actual bargain.The model estimates the hypothetical bargain, while judges retain responsibility for evaluating its significance.
- Perils and Boundaries: The approach faces limits involving model reliability, consumer consent, and one-off deals outside majoritarian linguistic communities.The authors identify hallucination, prompt sensitivity, sycophancy, randomness, and overconfidence as reliability concerns, while noting special problems for consumer and atypical contracts.
IV. CONCLUSION
The paper argues that generative gap filling can make contractual interpretation more precise, while AI-authored contracts raise unresolved questions about human intent and legal doctrine.
- Choice of Model clauses, adversarial presentation, and disclosure can discipline AI-assisted contractual interpretation.The authors present these tools as ways to reduce AI pathologies and make opinions narrower and more contestable.
- AI-assisted precision may reduce interpretation disputes, leaving courts to resolve the harder normative question of whether to honor what parties would have said.The remaining disputes would concern normative interpretation rather than merely recovering intended meaning.
- Contracts drafted by AI agents challenge the article’s assumption that a human bargain generated the information linking visible and hidden terms.The paper’s empirical foundation is a claim about how human bargains are made, not about text in the abstract.
- Current legal doctrines lack a clear way to attribute clause-level meaning to a principal who never read, drafted, or contemplated an AI-generated clause.The authors suggest upstream instructions and model prompts might become relevant, but expressly leave the doctrinal problem unsolved.
- For human-drafted agreements, the large existing stock of contracts means generative gap filling remains applicable for the foreseeable future.The authors distinguish this continuing domain from the stranger problems posed by machine-written contracts.