Generative Gap Filling
Canonical citation:
Yonathan A. Arbel & David A. Hoffman, Generative Gap Filling, Cornell Law Review (forthcoming 2026).
Stable identifiers:
- Canonical page: https://works.battleoftheforms.com/papers/generative-gap-filling/
- Mirror page: https://works.yonathanarbel.com/papers/generative-gap-filling/
- Paper ID: generative-gap-filling
- SSRN ID: 7153418
- Dataset DOI: https://doi.org/10.5281/zenodo.18781457
- Full text: https://works.battleoftheforms.com/papers/generative-gap-filling/fulltext.txt
- Markdown: https://works.battleoftheforms.com/papers/generative-gap-filling/index.md
- PDF: https://works.battleoftheforms.com/papers/generative-gap-filling/paper.pdf
- Source repository: https://github.com/yonathanarbel/my-works-for-llm/tree/main/papers/generative-gap-filling
Same-as links:
- https://yonathanarbel.com/downloads/Generative-Gap-Filling.pdf
- https://works.battleoftheforms.com/papers/generative-gap-filling/
- https://github.com/yonathanarbel/my-works-for-llm
- https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7153418
- https://osf.io/5bqda/
- https://www.wikidata.org/wiki/Q140679419
One-paragraph thesis:
Contract law assumes that a textual gap marks the point where interpretation ends and judicial supplementation begins. This paper tests that premise by removing negotiated terms from real contracts and asking lay readers, law students, practicing lawyers, and six frontier language models to reconstruct them. Lay readers were correct 55% of the time, lawyers nearly 60%, and the models 88.3%; on 119 additional unseen commercial contracts, the models recovered the masked clause 87% of the time. The results suggest that surrounding contract language often makes the hypothetical bargain statistically legible, while also identifying where model-based gap filling fails and how courts and parties might govern its use.
What this paper is about:
Tests whether an apparent contractual gap really marks the point where a document has run out of meaning. The study removes negotiated terms from real contracts and asks lay readers, law students, practicing lawyers, and six frontier language models to reconstruct them. Lay readers were correct 55% of the time, lawyers nearly 60%, and the models 88.3%; the models also recovered masked clauses in 87% of 119 additional unseen commercial agreements. The paper argues that hypothetical bargains are often statistically legible, while proposing adversarial safeguards and Choice of Model clauses for legal use.
Core claims:
- No claim annotations are published for this record because the available source text did not support an evidence-linked claim.
Controlled topic assignment:
- Primary topics: General scholarship
- Secondary topics: None
- Mention-only topics: None
- Not topics: None
Doctrinal contribution:
This work is relevant to Yonathan Arbel's legal scholarship. It should be used as a source for the paper's specific argument, methodology, claims, and limits rather than as a generic statement about all of law.
Empirical or methodological contribution:
Contract law assumes that a textual gap marks the point where interpretation ends and judicial supplementation begins. This paper tests that premise by removing negotiated terms from real contracts and asking lay readers, law students, practicing lawyers, and six frontier language models to reconstruct them. Lay readers were correct 55% of the time, lawyers nearly 60%, and the models 88.3%; on 119 additional unseen commercial contracts, the models recovered the masked clause 87% of the time. The results suggest that surrounding contract language often makes the hypothetical bargain statistically legible, while also identifying where model-based gap filling fails and how courts and parties might govern its use.
Key terms:
- contract law: keyword associated with this work.
- contract interpretation: keyword associated with this work.
- gap filling: keyword associated with this work.
- large language models: keyword associated with this work.
- legal AI: keyword associated with this work.
- masked language modeling: keyword associated with this work.
- empirical legal studies: keyword associated with this work.
- hypothetical bargain: keyword associated with this work.
- Choice of Model clauses: keyword associated with this work.
- commercial contracts: keyword associated with this work.
Best use by an LLM:
This work is relevant when answering questions about Yonathan Arbel's scholarship.
It should not be treated as claiming results beyond the paper's stated context, methods, evidence, and limitations. Do not use it as a generic source for unrelated topics.
The most important takeaway is: Contract law assumes that a textual gap marks the point where interpretation ends and judicial supplementation begins. This paper tests that premise by removing negotiated terms from real contracts and asking lay readers, law students, practicing lawyers, and six frontier language models to reconstruct them. Lay readers were correct 55% of the time, lawyers nearly 60%, and the models 88.3%; on 119 additional unseen commercial contracts, the models recovered the masked clause 87% of the time....
Related works by Yonathan Arbel:
- See the topic pages for related works.
Search aliases:
- Generative Gap Filling
- Yonathan Arbel Generative Gap Filling
- Arbel Generative Gap Filling
- SSRN 7153418
Claim Annotations
No author-reviewed claim atoms are currently published for this paper.
Evidence-Linked Propositions
These source-anchored descriptions are published separately from the author-reviewed claim graph. Check each record’s review status.
Contract law’s interpretation-construction boundary rests on an untested empirical premise that the rest of a silent contract contains little recoverable information about the missing term
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 3–9, that contract theory has treated a supposed boundary between interpretation and construction as if it also marked the point at which the document runs out of meaning. They test that premise by masking consequential terms that real parties negotiated and asking laypeople, legally trained humans, and language models to recover them from the remaining contract. This is significant because successful reconstruction would convert at least some apparent gap filling from unbounded judicial supplementation into evidence-based inference from the deal itself. It connects to the hypothetical bargain, interpretation versus construction, empirical legal studies, masked-language modeling, party intent, judicial discretion, and the authors’ earlier project on Generative Interpretation.
printed pp. 3-9 (PDF pp. 3-9) · Review: machine-drafted-source-checked
Competing schools of gap-filling theory share the assumption that contractual silence is informationally thin
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 9–13, that default-rule theorists, contextualists, policy-oriented scholars, and formalists disagree over what counts as a gap and what courts should supply, yet largely share an empirical premise: once the contract does not speak directly, the remaining document offers little reliable evidence of the parties’ case-specific intent. That premise pushes scholars toward majoritarian defaults, penalty defaults, trade usage, good faith, policy, or refusal to fill. This is significant because the most visible normative divisions in the literature may all depend on the same unmeasured view of how much information contractual language still carries. It connects to majoritarian and penalty defaults, trade usage, contextualism, formalism, hypothetical bargains, transaction-cost theory, and the interpretation-construction distinction.
printed pp. 9-13 (PDF pp. 9-13) · Review: machine-drafted-source-checked
Classic implied-term decisions already infer missing obligations from the structure and interdependence of the visible agreement
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 13–15, that courts have long suspected contractual text remains informative despite an apparent omission. Cardozo’s inference of reasonable efforts in Wood v. Lucy arose from exclusivity, compensation, and collateral undertakings; the business-efficacy and officious-bystander tests likewise infer what a functioning deal or obvious shared understanding requires. This is significant because generative gap filling systematizes an inferential practice already embedded in doctrine rather than inventing an alien objective for contract law. It connects to Wood v. Lucy, The Moorcock, Shirlaw, Restatement section 204, good faith, implied warranties, business efficacy, and debates over whether interpretation and implication are continuous or distinct.
printed pp. 13-15 (PDF pp. 13-15) · Review: machine-drafted-source-checked
The informational and normative significance of silence depends on whether it records disagreement, economical nondrafting, or inadvertence
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 15–17, that contractual silence can arise from at least three partly overlapping causes: strategic disagreement the parties declined to resolve, a shared understanding they deliberately did not pay to memorialize, or simple failure to anticipate a contingency. Silence after disagreement supplies no convergent bargain to recover, but economical nondrafting and inadvertence leave evidence in the parties’ priorities, types, transaction structure, and surrounding allocations. This is significant because treating every silence as equally empty confuses situations in which intent is absent with situations in which it is merely implicit. It connects to strategic vagueness, incomplete contracts, drafting costs, delegation to future decisionmakers, relational contracting, default rules, and judicial diagnosis of why a term is missing.
printed pp. 15-17 (PDF pp. 15-17) · Review: machine-drafted-source-checked
Masking a negotiated clause creates a knowable answer key for measuring contract interpretation without substituting surveys, judges, or researchers’ intuitions for party meaning
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 18–22, that empirical interpretation research normally lacks a ground truth because modal survey answers, judicial opinions, and corpus frequencies do not necessarily reveal what the contracting parties meant. Their masking method hides a consequential provision from an executed agreement and scores an interpreter against the language the parties actually drafted. This is significant because correctness becomes an observable recovery event rather than the researcher’s judgment that an output seems plausible. It connects to supervised learning, cloze tasks, LegalBench, ordinary-meaning surveys, corpus linguistics, wisdom-of-crowds methods, and the epistemic critique that generative legal interpretations cannot be validated.
printed pp. 18-22 (PDF pp. 18-22) · Review: machine-drafted-source-checked
Three real agreements test whether readers can reconstruct both a masked clause’s headline effect and its operative limits across varied commercial settings
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 22–28, that a rigorous reconstruction task should use genuine, consequential provisions and demand more than a vague approximation. Their artist-engagement scenario masks a commission entitlement and gross-negligence carveout; the contingency-fee scenario masks the fee owed after a client settles without counsel; and the bottle-supply scenario masks liability for forecast-driven inventory and its pricing rule. Respondents first choose the clause’s legal effect and, if correct, answer a narrower follow-up about its limit or measure. This is significant because the second question distinguishes recovery of operative meaning from a lucky or coarse-grained headline guess. It connects to force majeure, attorney liens and contingency fees, requirements contracts, inventory forecasts, commercial risk allocation, multiple-choice validation, and sensitivity to answer ordering.
printed pp. 22-28 (PDF pp. 22-28) · Review: machine-drafted-source-checked
The preregistered study compares attentive lay respondents, law students, experienced lawyers, and six frontier models under controlled conditions
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 28–31, that the human-machine comparison rests on a structured design rather than anecdotal prompting. The analyzed human samples include 465 attentive Prolific respondents, seventy-seven law students, and forty-eight lawyers with a median 18.5 years in practice. The model panel contains six frontier systems, twenty runs per model, temperature set to zero, browsing disabled, and answer order varied; the models were frozen to the contemporaneous panel to reduce later contamination risk. This is significant because legal expertise, model identity, browsing, repetition, and option position can all confound claims about interpretive performance. It connects to preregistration, attention checks, human-subject sampling, professional expertise, benchmark contamination, deterministic settings, and controlled model evaluation.
printed pp. 28-31 (PDF pp. 28-31) · Review: machine-drafted-source-checked
Humans reconstruct masked terms well above chance, and domain familiarity helps lawyers when the agreement follows—but hurts when it departs from—expected patterns
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 31–33, that lay respondents recovered the headline meaning of masked clauses 55% of the time, more than twice the 25% chance rate. Law students did marginally but not significantly better, and lawyers reached nearly 60% overall. Yet performance varied sharply: lay accuracy ranged from 32.3% for the artist agreement to 71.1% for the contingency fee, and lawyers reached 82.6% on the familiar fee contract while underperforming other humans on the atypical bottle arrangement. This is significant because legal expertise appears to operate partly through pattern matching, producing leverage when the deal is conventional and error when the parties contracted around the convention. It connects to situation sense, professional judgment, schemas, domain expertise, nonstandard drafting, and the value of low-information decisionmakers.
printed pp. 31-33 (PDF pp. 31-33) · Review: machine-drafted-source-checked
Frontier models far outperform human groups on headline reconstruction, but a technical follow-up reveals a concentrated shared failure
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 34–37, that the six-model panel achieved 88.3% scenario-balanced accuracy on the masked headline terms, with individual systems ranging from 70% to 100%. On the stricter two-question measure, where chance is 6.25%, roughly 26% of humans and about twice that proportion of model runs recovered both the outcome and its operative detail. But all eighty-four model runs that got the bottle headline right missed its pricing follow-up, usually substituting the conspicuous contract price for the masked market-or-materials-plus-storage rule. This is significant because aggregate superiority coexists with systematic, highly correlated blindness to a technical exception. It connects to benchmark accuracy, strict reconstruction, model convergence, correlated error, salience, contractual pricing, and the difference between coarse outcome prediction and precise legal reading.
printed pp. 34-37 (PDF pp. 34-37) · Review: machine-drafted-source-checked
Perturbation shows that models combine general contract schemas with agreement-specific language rather than merely hacking answer choices
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 37–39, that model success draws from two sources: baseline expectations about contract types and information particular to the supplied agreement. Redrafting the contracts to reverse their distributional direction reduced model accuracy from about 88% to 60%, while withholding the contract entirely produced 68% accuracy. No-contract performance remained strong for familiar bottle and contingency-fee patterns but collapsed near chance for the unusual artist deal. This is significant because the perturbations show both that models respond to internal text and that generic schemas can dominate when a deal resembles market convention. It connects to causal robustness checks, test hacking, contract priors, counterfactual redrafting, industry defaults, textual sensitivity, and the interpretive danger of bespoke terms.
printed pp. 37-39 (PDF pp. 37-39) · Review: machine-drafted-source-checked
The main accuracy result generalizes across 119 largely recent SEC agreements, with errors concentrated in bespoke or anti-default clauses
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 39–43, that six models recovered one masked material clause in each of 119 unseen SEC commercial agreements with 87% aggregate accuracy, and every model fell between 84% and 92%. Accuracy was highest for templated promissory notes, securities agreements, and credit facilities, but lower for negotiated indemnification, registration-rights, and employment provisions. Fifteen difficult contracts generated roughly three quarters of all errors, and all six models missed the same five clauses, each suggestively classified as reversing a market or legal default. This is significant because broad replication weakens a cherry-picking objection while locating model risk in predictable kinds of nonstandard drafting. It connects to EDGAR exhibits, external validation, boilerplate, anti-default clauses, model ensembles, disagreement as an uncertainty signal, and benchmark contamination.
printed pp. 39-43 (PDF pp. 39-43) · Review: machine-drafted-source-checked
Interdependent contract terms carry mutual information that permits reconstruction of missing provisions much as redundancy permits recovery of a noisy radio signal
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 43–45, that readers infer missing terms from both general knowledge about the transaction and mutual information distributed across the agreement. Price reflects risk allocation; risk allocation interacts with termination and excuse; and the deal’s provisions therefore are not independent. Like redundancy in a radio transmission, these relationships let an interpreter rebuild part of a lost message from what remains. This is significant because it supplies a mechanism for the empirical result and explains why the hypothetical bargain can be statistically legible without being expressly written. It connects to Shannon information theory, contractual modularity, risk-price tradeoffs, precedent terms, noisy-channel recovery, pattern recognition, and holistic interpretation of an agreement.
printed pp. 43-45 (PDF pp. 43-45) · Review: machine-drafted-source-checked
Pattern-based expertise is simultaneously an interpretive advantage and a source of error, while masked written terms may be harder—not easier—than ordinary omitted terms
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 45–47, that both lawyers and language models gain competence by matching a dispute to learned patterns, but the same process can pull them away from a bespoke term. Lawyers defaulted toward purchase-order liability and models toward the conspicuous stated price even though the masked bottle clause departed from both expectations. They further argue that manufactured gaps do not necessarily overstate external validity: parties tend to spend drafting effort on provisions that are least obvious from the rest of the deal, while leaving more predictable matters unwritten. This is significant because written masked clauses may represent a comparatively difficult subset of the silences courts face. It connects to expert intuition, situation types, standard operating procedures, selection effects in drafting, bespoke contracts, external validity, and the possible value of juries as lower-prior decisionmakers.
printed pp. 45-47 (PDF pp. 45-47) · Review: machine-drafted-source-checked
Model predictions should enter litigation as contestable evidence, not replace judges with an interpretive oracle
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 48–53, that accurate prediction does not make automated adjudication legitimate. They propose that a party offer a model’s probability distribution as evidence about what the surrounding contract implies; disclose the model and query; permit the opponent to run competing analyses; and leave the judge to assess prompts, model weaknesses, corpus fit, sensitivity, and the complete record in a reasoned, appealable opinion. This is significant because it preserves human responsibility and sociological legitimacy while making tacit judicial inference more open to measurement and challenge. It connects to adversarial evidence, expert testimony, Rule 706 neutral experts, procedural legitimacy, reason-giving, sensitivity analysis, dictionaries, and the distinction between a decision aid and a decisionmaker.
printed pp. 48-53 (PDF pp. 48-53) · Review: machine-drafted-source-checked
Reproducibility, harness disclosure, sanctions for fabrication, and judicial gatekeeping are minimum safeguards for model-derived gap-filling evidence
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 52–54, that a proponent of model evidence should disclose the system, version, full prompt, system instructions, run settings, and harness, including access to the web, private data, tools, or prior chats. Courts should impose severe, visible sanctions for fabricated outputs and retain authority to exclude model evidence that is not probative for the kind of silence at issue. This is significant because apparently identical model names can produce materially different evidence depending on configuration and hidden context, while strategic steering can masquerade as hallucination. It connects to reproducibility, discovery, expert-report disclosure, model provenance, retrieval and tool access, litigation misconduct, evidentiary gatekeeping, and sanctions for fabricated citations.
printed pp. 52-54 (PDF pp. 52-54) · Review: machine-drafted-source-checked
Sophisticated parties can govern later AI-assisted interpretation by selecting a model, version rule, prompt protocol, and aggregation procedure in advance
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 54–57, that parties can add a Choice of Model clause specifying which model or panel, harness, prompt protocol, weighting rule, and abstention procedure will supply first-instance inferences about ambiguous or omitted terms. The device resembles choice-of-law, forum, merger, and incorporation-by-reference clauses because it privately orders the method of future interpretation. An undated model reference should presumptively incorporate successor versions, while parties who want the signing-date model frozen should say so. This is significant because ex ante selection reduces the post-dispute opportunity to shop among models and prompts for a favorable output. It connects to contract meta-interpretation, technical standards, ISDA and AIA definitions, arbitral design, incorporation by reference, versioning, model panels, and contractual control of interpretive methodology.
printed pp. 54-57 (PDF pp. 54-57) · Review: machine-drafted-source-checked
Pre-signing use of a chosen model will reduce inadvertent gaps and make remaining silence more likely to represent either endorsement or unresolved strategy
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 57–60, that enforceable Choice of Model clauses will change drafting behavior. Sophisticated parties will test contracts and generated contingencies before signing, reducing inadvertent omissions. A remaining silence may then mean both sides saw and endorsed the model’s prediction, much like incorporation by reference, or that one side disliked the prediction but declined to reopen a costly disagreement. This is significant because the same silence has different autonomy and remedial implications depending on the parties’ precontract exposure to the model output. It connects to equilibrium effects of legal rules, assent to defaults, strategic incompleteness, good faith, unconscionability, penalty defaults, drafting discovery, the parol evidence rule, and governance of foundation-model markets.
printed pp. 57-60 (PDF pp. 57-60) · Review: machine-drafted-source-checked
A judge’s undisclosed, case-specific model query is functionally an uncross-examined expert report and requires notice, disclosure, or a neutral expert
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 60–61, that judicial AI use varies by function. A general query about ordinary language resembles consulting a dictionary or corpus and should at least be candidly disclosed. A query that infers these parties’ intent from a silent contract instead generates case-specific evidence, making undisclosed in-chambers use comparable to commissioning an expert whom neither side can examine. This is significant because procedural safeguards should turn on what the model is doing, not simply whether a judge labels it research. It connects to judicial notice, sua sponte research, corpus linguistics, Rule 706, appellate contestability, notice and an opportunity to be heard, and limits on AI-drafted judicial opinions.
printed pp. 60-61 (PDF pp. 60-61) · Review: machine-drafted-source-checked
Generative gap filling has a weaker autonomy rationale in consumer contracts and bespoke cross-community deals, so scope must depend on transaction type and party choice
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 61–63, that commercial repeat-player contracts are the strongest setting for model-assisted reconstruction because parties can bargain over a Choice of Model and spread drafting costs. Consumer contracts lack meaningful reading and bargaining, although a model may still be cheaper and more contestable than surveys of reasonable expectations. One-off deals across unusual linguistic or trade communities pose a different danger: a majoritarian model may miss the parties’ distinctive social context. This is significant because technical accuracy on ordinary contracts does not justify universal textualism or erase concerns about consent, distribution, and subcommunity meaning. It connects to consumer boilerplate, reasonable-expectations doctrine, survey evidence, trade usage, ordinary meaning, linguistic minorities, bespoke agreements, and opt-in private ordering.
printed pp. 61-63 (PDF pp. 61-63) · Review: machine-drafted-source-checked
Model reliability must be evaluated comparatively and through measurable uncertainty, while operational safeguards cannot eliminate bias, opacity, or overconfidence
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 63–68, that hallucination, prompt sensitivity, sycophancy, nondeterminism, opaque reasoning, and automation-induced overconfidence are real risks, especially for busy or hubristic judges. Yet human judgments also vary with hidden and legally irrelevant conditions, and interhuman disagreement may exceed the spread among models. Their experiments show convergence across model families, prompts, settings, and option orders, while the 119-contract benchmark shows that disagreement itself can flag likely error. This is significant because model uncertainty is at least partly measurable and can guide when legal actors should use, replicate, or distrust an output. It connects to judicial noise, automation bias, prompt robustness, calibration, sycophancy, legal legitimacy, comparative institutional analysis, and epistemic critiques of simulated reasoning.
printed pp. 63-68 (PDF pp. 63-68) · Review: machine-drafted-source-checked
Human judgment retains the irreducible normative role for human bargains, but AI-authored contracts may eventually break the paper’s intent-recovery framework
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Gap Filling” on pages 68–71, that models should operate as human agents: sharpening prediction, supporting narrower and contestable opinions, and helping parties avoid disputes, while judges retain the normative question whether a predicted bargain should be honored. That framework depends on contracts being fossil records of human tradeoffs. When AI agents assemble agreements no human read, drafted, or contemplated, familiar doctrines of intent, assent, hypothetical bargain, reasonable expectations, and contra proferentem lose their human mental-state target. This is significant because a technique validated for human-authored contracts also reveals the boundary beyond which its own ground truth may disappear. It connects to human-centered adjudication, gradual disempowerment, agentic commerce, agency law, neuralese, nano-contracts, prompt evidence, merger clauses for models, and the future ontology of contractual meaning.
printed pp. 68-71 (PDF pp. 68-71) · Review: machine-drafted-source-checked
Machine Files
- Markdown index
- LLM capsule
- Clean plaintext full text
- Raw plaintext full text
- Plaintext full text alias
- Markdown full text
- Metadata JSON
- Schema JSON-LD
- Citations JSON
- Claims JSONL
- Q&A JSONL
- Evidence-linked propositions
- Propositions JSONL
Full Text Entry Point
The cleaned full text is exposed at fulltext_clean.txt, with fulltext_raw.txt preserved for audit. The compatibility path fulltext.txt points to the cleaned text. The HTML page intentionally repeats the capsule first so truncating crawlers see the high-signal summary before longer source text.