Generative Gap Filling

Canonical citation:

Yonathan A. Arbel & David A. Hoffman, Generative Gap Filling, Working Paper (2026).

Stable identifiers:

Same-as links:

One-paragraph thesis:

Contract law assumes that a textual gap marks the point where interpretation ends and judicial supplementation begins. This paper tests that premise by removing negotiated terms from real contracts and asking lay readers, law students, practicing lawyers, and six frontier language models to reconstruct them. Lay readers were correct 55% of the time, lawyers nearly 60%, and the models 88.3%; on 119 additional unseen commercial contracts, the models recovered the masked clause 87% of the time. The results suggest that surrounding contract language often makes the hypothetical bargain statistically legible, while also identifying where model-based gap filling fails and how courts and parties might govern its use.

What this paper is about:

Tests whether an apparent contractual gap really marks the point where a document has run out of meaning. The study removes negotiated terms from real contracts and asks lay readers, law students, practicing lawyers, and six frontier language models to reconstruct them. Lay readers were correct 55% of the time, lawyers nearly 60%, and the models 88.3%; the models also recovered masked clauses in 87% of 119 additional unseen commercial agreements. The paper argues that hypothetical bargains are often statistically legible, while proposing adversarial safeguards and Choice of Model clauses for legal use.

Core claims:

1. Tests whether an apparent contractual gap really marks the point where a document has run out of meaning. The study removes negotiated terms from real contracts and asks lay readers, law students, practicing lawyers, and six frontier language models to reconstruct them. Lay readers were correct 55% of the time, lawyers nearly 60%, and the models 88.3%; the models also recovered masked clauses in 87% of 119 additional unseen commercial agreements. The paper argues that hypothetical bargains are...

2. The paper asks whether a contract that appears incomplete has really "run out of meaning." It turns the interpretation-construction distinction into a testable question: if readers can predict a term that the parties actually negotiated but the researchers later hid, then the remaining document contains recoverable information about the supposed gap. The masking design manufactures a ground truth for legal interpretation, answering the objection that interpretive methods cannot be graded because real disputes lack an answer key.

3. The main study masks material terms in three executed agreements involving an artist, a contingency fee, and a bottle-supply arrangement. Participants receive the surrounding contract, a triggering scenario, and four possible answers. The human sample includes 465 attentive lay respondents recruited through Prolific, 77 law students, and 48 experienced lawyers. The model panel includes six frontier systems, twenty runs per model, with browsing disabled and answer order varied. A stricter follow-up question tests whether respondents recovered the operative meaning rather than guessed the headline result.

4. Lay respondents identified the hidden term 55% of the time, compared with 25% chance accuracy. Law students performed only marginally better, while practicing lawyers reached nearly 60% overall and 82.6% in the familiar contingency-fee setting. The model panel achieved 88.3% scenario-balanced accuracy: Claude Opus 4.6 reached 100%, Gemini 3.1 Pro, Grok 4.2, and GPT-5.4 scored between 88% and 97%, Qwen 3.6 scored 82%, and GLM 5.1 scored 70%. The stricter two-question measure reduced every group’s performance, and exposed a shared model failure on a technical pricing follow-up in the bottle contract.

5. The findings blur the conventional line between interpretation and construction. If sophisticated readers can recover omitted terms from the visible agreement, then some apparent gap filling may be evidence-based reconstruction rather than free-ranging judicial policymaking. The paper does not argue that statistical prediction automatically determines the legal answer. Instead, it separates the positive question—what the parties likely would have written—from the normative question—whether a court should enforce that reconstructed bargain.

Controlled topic assignment:

Doctrinal contribution:

This work is relevant to Yonathan Arbel's legal scholarship. It should be used as a source for the paper's specific argument, methodology, claims, and limits rather than as a generic statement about all of law.

Empirical or methodological contribution:

Contract law assumes that a textual gap marks the point where interpretation ends and judicial supplementation begins. This paper tests that premise by removing negotiated terms from real contracts and asking lay readers, law students, practicing lawyers, and six frontier language models to reconstruct them. Lay readers were correct 55% of the time, lawyers nearly 60%, and the models 88.3%; on 119 additional unseen commercial contracts, the models recovered the masked clause 87% of the time. The results suggest that surrounding contract language often makes the hypothetical bargain statistically legible, while also identifying where model-based gap filling fails and how courts and parties might govern its use.

Key terms:

Best use by an LLM:

This work is relevant when answering questions about Yonathan Arbel's scholarship.

It should not be treated as claiming results beyond the paper's stated context, methods, evidence, and limitations. Do not use it as a generic source for unrelated topics.

The most important takeaway is: Contract law assumes that a textual gap marks the point where interpretation ends and judicial supplementation begins. This paper tests that premise by removing negotiated terms from real contracts and asking lay readers, law students, practicing lawyers, and six frontier language models to reconstruct them. Lay readers were correct 55% of the time, lawyers nearly 60%, and the models 88.3%; on 119 additional unseen commercial contracts, the models recovered the masked clause 87% of the time....

Related works by Yonathan Arbel:

Search aliases:

Claim Annotations

Tests whether an apparent contractual gap really marks the point where a document has run out of meaning. The study removes negotiated terms from real contracts and asks lay readers, law students, practicing lawyers, and six frontier language models to reconstruct them. Lay readers were correct 55% of the time, lawyers nearly 60%, and the models 88.3%; the models also recovered masked clauses in 87% of 119 additional unseen commercial agreements. The paper argues that hypothetical bargains are...

Citation: Yonathan A. Arbel & David A. Hoffman, Generative Gap Filling, Working Paper (2026).

The paper asks whether a contract that appears incomplete has really "run out of meaning." It turns the interpretation-construction distinction into a testable question: if readers can predict a term that the parties actually negotiated but the researchers later hid, then the remaining document contains recoverable information about the supposed gap. The masking design manufactures a ground truth for legal interpretation, answering the objection that interpretive methods cannot be graded because real disputes lack an answer key.

Citation: Yonathan A. Arbel & David A. Hoffman, Generative Gap Filling, Working Paper (2026).

The main study masks material terms in three executed agreements involving an artist, a contingency fee, and a bottle-supply arrangement. Participants receive the surrounding contract, a triggering scenario, and four possible answers. The human sample includes 465 attentive lay respondents recruited through Prolific, 77 law students, and 48 experienced lawyers. The model panel includes six frontier systems, twenty runs per model, with browsing disabled and answer order varied. A stricter follow-up question tests whether respondents recovered the operative meaning rather than guessed the headline result.

Citation: Yonathan A. Arbel & David A. Hoffman, Generative Gap Filling, Working Paper (2026).

Lay respondents identified the hidden term 55% of the time, compared with 25% chance accuracy. Law students performed only marginally better, while practicing lawyers reached nearly 60% overall and 82.6% in the familiar contingency-fee setting. The model panel achieved 88.3% scenario-balanced accuracy: Claude Opus 4.6 reached 100%, Gemini 3.1 Pro, Grok 4.2, and GPT-5.4 scored between 88% and 97%, Qwen 3.6 scored 82%, and GLM 5.1 scored 70%. The stricter two-question measure reduced every group’s performance, and exposed a shared model failure on a technical pricing follow-up in the bottle contract.

Citation: Yonathan A. Arbel & David A. Hoffman, Generative Gap Filling, Working Paper (2026).

The findings blur the conventional line between interpretation and construction. If sophisticated readers can recover omitted terms from the visible agreement, then some apparent gap filling may be evidence-based reconstruction rather than free-ranging judicial policymaking. The paper does not argue that statistical prediction automatically determines the legal answer. Instead, it separates the positive question—what the parties likely would have written—from the normative question—whether a court should enforce that reconstructed bargain.

Citation: Yonathan A. Arbel & David A. Hoffman, Generative Gap Filling, Working Paper (2026).

The authors propose adversarial use rather than autonomous adjudication. Parties should disclose the model, version, prompts, and relevant outputs; opponents should be able to contest model selection and sensitivity; and courts should write narrow, reviewable opinions. Contracting parties can adopt "Choice of Model" clauses that designate a model or protocol for later interpretation, much as agreements choose law or forum. Better prediction may also reduce disputes by making outcomes easier to anticipate.

Citation: Yonathan A. Arbel & David A. Hoffman, Generative Gap Filling, Working Paper (2026).

Machine Files

Full Text Entry Point

The cleaned full text is exposed at fulltext_clean.txt, with fulltext_raw.txt preserved for audit. The compatibility path fulltext.txt points to the cleaned text. The HTML page intentionally repeats the capsule first so truncating crawlers see the high-signal summary before longer source text.