{"claim_id": "generative-gap-filling-001", "claim": "Tests whether an apparent contractual gap really marks the point where a document has run out of meaning. The study removes negotiated terms from real contracts and asks lay readers, law students, practicing lawyers, and six frontier language models to reconstruct them. Lay readers were correct 55% of the time, lawyers nearly 60%, and the models 88.3%; the models also recovered masked clauses in 87% of 119 additional unseen commercial agreements. The paper argues that hypothetical bargains are...", "paper_id": "generative-gap-filling", "paper_title": "Generative Gap Filling", "claim_type": "core_thesis", "evidence_quote": "GENERATIVE GAP FILLING Yonathan A. Arbel David A. Hoffman Contract law polices a line between interpretation, the recovery of meaning a text already holds, and gap filling, the supply of terms the text lacks. The boundary rests on an unchecked premise: that courts need to fill gaps themselves because the document has run out of meaning. We tested it. Borrowing a masking design from machine learning, we took real contracts, hid a term the parties had negotiated, and asked three kinds of readers to predict what we had removed: ordinary people, legally trained ones, and several large language models. Human interpreters guessed right a bit more than half the time, doubling the rate of chance...", "evidence_page": null, "evidence_span": "GENERATIVE GAP FILLING Yonathan A. Arbel David A. Hoffman Contract law polices a line between interpretation, the recovery of meaning a text already holds, and gap filling, the supply of terms the text lacks. The boundary rests on an unchecked premise: that courts need to fill gaps themselves because the document has run out of meaning. We tested it. Borrowing a masking design from machine learning, we took real contracts, hid a term the parties had negotiated, and asked three kinds of readers to predict what we had removed: ordinary people, legally trained ones, and several large language models. Human interpreters guessed right a bit more than half the time, doubling the rate of chance...", "source_text_url": "https://works.battleoftheforms.com/papers/generative-gap-filling/fulltext_clean.txt", "canonical_url": "https://works.battleoftheforms.com/papers/generative-gap-filling/#claim-001", "citation": "Yonathan A. Arbel & David A. Hoffman, Generative Gap Filling, Working Paper (2026).", "topics": [], "secondary_topics": [], "human_reviewed": false, "confidence": "machine-linked", "limitations": "Machine-linked claim. Use the evidence quote and PDF before treating it as a quotation or as a complete statement of the paper's position."}
{"claim_id": "generative-gap-filling-002", "claim": "The paper asks whether a contract that appears incomplete has really \"run out of meaning.\" It turns the interpretation-construction distinction into a testable question: if readers can predict a term that the parties actually negotiated but the researchers later hid, then the remaining document contains recoverable information about the supposed gap. The masking design manufactures a ground truth for legal interpretation, answering the objection that interpretive methods cannot be graded because real disputes lack an answer key.", "paper_id": "generative-gap-filling", "paper_title": "Generative Gap Filling", "claim_type": "supporting_claim", "evidence_quote": "language models could help courts parse contract text.18 That argument drew many heated objections, one we half-conceded ourselves: we could show a model’s reading was sensible, but not that it was correct, because in a real dispute there is no answer key.19 One carefully-argued response put it starkly: “no experiment can determine whether a generative method yields correct results, because there is no accessible source of ground truth for legal meaning.”20 Here, by redacting a term that was in fact negotiated, priced, drafted, and signed, we can manufacture a ground truth and, for the first time, grade interpretative predictions against it.21 In our masking experiment, humans were...", "evidence_page": null, "evidence_span": "language models could help courts parse contract text.18 That argument drew many heated objections, one we half-conceded ourselves: we could show a model’s reading was sensible, but not that it was correct, because in a real dispute there is no answer key.19 One carefully-argued response put it starkly: “no experiment can determine whether a generative method yields correct results, because there is no accessible source of ground truth for legal meaning.”20 Here, by redacting a term that was in fact negotiated, priced, drafted, and signed, we can manufacture a ground truth and, for the first time, grade interpretative predictions against it.21 In our masking experiment, humans were...", "source_text_url": "https://works.battleoftheforms.com/papers/generative-gap-filling/fulltext_clean.txt", "canonical_url": "https://works.battleoftheforms.com/papers/generative-gap-filling/#claim-002", "citation": "Yonathan A. Arbel & David A. Hoffman, Generative Gap Filling, Working Paper (2026).", "topics": [], "secondary_topics": [], "human_reviewed": false, "confidence": "machine-linked", "limitations": "Machine-linked claim. Use the evidence quote and PDF before treating it as a quotation or as a complete statement of the paper's position."}
{"claim_id": "generative-gap-filling-003", "claim": "The main study masks material terms in three executed agreements involving an artist, a contingency fee, and a bottle-supply arrangement. Participants receive the surrounding contract, a triggering scenario, and four possible answers. The human sample includes 465 attentive lay respondents recruited through Prolific, 77 law students, and 48 experienced lawyers. The model panel includes six frontier systems, twenty runs per model, with browsing disabled and answer order varied. A stricter follow-up question tests whether respondents recovered the operative meaning rather than guessed the headline result.", "paper_id": "generative-gap-filling", "paper_title": "Generative Gap Filling", "claim_type": "supporting_claim", "evidence_quote": "For human interpreters, the artist and bottle follow up questions presented a significant challenge, while the contingency fee follow up was significantly easier. LLMs had a different nemesis. They breezed through the artist and contingency fee questions, but were simply unable to correctly answer the follow-up bottle scenario question: among the eighty-four runs that correctly answered the headline bottles question, none selected the correct follow-up answer; eighty-one chose option C and three chose option B. Recall that this option held that CKS was only entitled to the contract price, rather than current market price or raw materials and storage costs. Presumably, the models...", "evidence_page": null, "evidence_span": "For human interpreters, the artist and bottle follow up questions presented a significant challenge, while the contingency fee follow up was significantly easier. LLMs had a different nemesis. They breezed through the artist and contingency fee questions, but were simply unable to correctly answer the follow-up bottle scenario question: among the eighty-four runs that correctly answered the headline bottles question, none selected the correct follow-up answer; eighty-one chose option C and three chose option B. Recall that this option held that CKS was only entitled to the contract price, rather than current market price or raw materials and storage costs. Presumably, the models...", "source_text_url": "https://works.battleoftheforms.com/papers/generative-gap-filling/fulltext_clean.txt", "canonical_url": "https://works.battleoftheforms.com/papers/generative-gap-filling/#claim-003", "citation": "Yonathan A. Arbel & David A. Hoffman, Generative Gap Filling, Working Paper (2026).", "topics": [], "secondary_topics": [], "human_reviewed": false, "confidence": "machine-linked", "limitations": "Machine-linked claim. Use the evidence quote and PDF before treating it as a quotation or as a complete statement of the paper's position."}
{"claim_id": "generative-gap-filling-004", "claim": "Lay respondents identified the hidden term 55% of the time, compared with 25% chance accuracy. Law students performed only marginally better, while practicing lawyers reached nearly 60% overall and 82.6% in the familiar contingency-fee setting. The model panel achieved 88.3% scenario-balanced accuracy: Claude Opus 4.6 reached 100%, Gemini 3.1 Pro, Grok 4.2, and GPT-5.4 scored between 88% and 97%, Qwen 3.6 scored 82%, and GLM 5.1 scored 70%. The stricter two-question measure reduced every group’s performance, and exposed a shared model failure on a technical pricing follow-up in the bottle contract.", "paper_id": "generative-gap-filling", "paper_title": "Generative Gap Filling", "claim_type": "supporting_claim", "evidence_quote": "The LLM panel is variably capable. Opus 4.6 scored 100% on scenario-balanced unmasking; Gemini 3.1 Pro, Grok 4.2, and GPT 5.4 ranged between 88% and 97%; Qwen 3.6 reached 82%; and GLM 5.1 trailed the panel at 70.0%, still well above the human rate. For context, the two trailing models are generally thought to be the weakest of the bunch.104 The headline measure asks whether the respondent identified the correct meaning of the redacted provision. We probed that result with a follow-up question designed to test a more specific implication of the same hidden language. That is, did the respondents get the right answer for the right reasons?", "evidence_page": null, "evidence_span": "The LLM panel is variably capable. Opus 4.6 scored 100% on scenario-balanced unmasking; Gemini 3.1 Pro, Grok 4.2, and GPT 5.4 ranged between 88% and 97%; Qwen 3.6 reached 82%; and GLM 5.1 trailed the panel at 70.0%, still well above the human rate. For context, the two trailing models are generally thought to be the weakest of the bunch.104 The headline measure asks whether the respondent identified the correct meaning of the redacted provision. We probed that result with a follow-up question designed to test a more specific implication of the same hidden language. That is, did the respondents get the right answer for the right reasons?", "source_text_url": "https://works.battleoftheforms.com/papers/generative-gap-filling/fulltext_clean.txt", "canonical_url": "https://works.battleoftheforms.com/papers/generative-gap-filling/#claim-004", "citation": "Yonathan A. Arbel & David A. Hoffman, Generative Gap Filling, Working Paper (2026).", "topics": [], "secondary_topics": [], "human_reviewed": false, "confidence": "machine-linked", "limitations": "Machine-linked claim. Use the evidence quote and PDF before treating it as a quotation or as a complete statement of the paper's position."}
{"claim_id": "generative-gap-filling-005", "claim": "The findings blur the conventional line between interpretation and construction. If sophisticated readers can recover omitted terms from the visible agreement, then some apparent gap filling may be evidence-based reconstruction rather than free-ranging judicial policymaking. The paper does not argue that statistical prediction automatically determines the legal answer. Instead, it separates the positive question—what the parties likely would have written—from the normative question—whether a court should enforce that reconstructed bargain.", "paper_id": "generative-gap-filling", "paper_title": "Generative Gap Filling", "claim_type": "supporting_claim", "evidence_quote": "Third, sanctions for fabricated outputs must be severe and visible, because the integrity of the entire mechanism depends on it. Recent episodes of lawyers submitting fake citations are a preview of what not doing this looks like.146 We use fabricated rather than hallucinated for a particular reason. There are subtle ways to steer models towards desired ends, and lawyer ingenuity knows no end. Litigants should be wary of risks and have the benefit of judicial penalties when opposing parties manipulate the tools in underhanded ways – much like the case of bribing an expert. Fourth, courts must retain a robust gatekeeping role for the threshold question of whether model output is probative...", "evidence_page": null, "evidence_span": "Third, sanctions for fabricated outputs must be severe and visible, because the integrity of the entire mechanism depends on it. Recent episodes of lawyers submitting fake citations are a preview of what not doing this looks like.146 We use fabricated rather than hallucinated for a particular reason. There are subtle ways to steer models towards desired ends, and lawyer ingenuity knows no end. Litigants should be wary of risks and have the benefit of judicial penalties when opposing parties manipulate the tools in underhanded ways – much like the case of bribing an expert. Fourth, courts must retain a robust gatekeeping role for the threshold question of whether model output is probative...", "source_text_url": "https://works.battleoftheforms.com/papers/generative-gap-filling/fulltext_clean.txt", "canonical_url": "https://works.battleoftheforms.com/papers/generative-gap-filling/#claim-005", "citation": "Yonathan A. Arbel & David A. Hoffman, Generative Gap Filling, Working Paper (2026).", "topics": [], "secondary_topics": [], "human_reviewed": false, "confidence": "machine-linked", "limitations": "Machine-linked claim. Use the evidence quote and PDF before treating it as a quotation or as a complete statement of the paper's position."}
{"claim_id": "generative-gap-filling-006", "claim": "The authors propose adversarial use rather than autonomous adjudication. Parties should disclose the model, version, prompts, and relevant outputs; opponents should be able to contest model selection and sensitivity; and courts should write narrow, reviewable opinions. Contracting parties can adopt \"Choice of Model\" clauses that designate a model or protocol for later interpretation, much as agreements choose law or forum. Better prediction may also reduce disputes by making outcomes easier to anticipate.", "paper_id": "generative-gap-filling", "paper_title": "Generative Gap Filling", "claim_type": "supporting_claim", "evidence_quote": "Third, sanctions for fabricated outputs must be severe and visible, because the integrity of the entire mechanism depends on it. Recent episodes of lawyers submitting fake citations are a preview of what not doing this looks like.146 We use fabricated rather than hallucinated for a particular reason. There are subtle ways to steer models towards desired ends, and lawyer ingenuity knows no end. Litigants should be wary of risks and have the benefit of judicial penalties when opposing parties manipulate the tools in underhanded ways – much like the case of bribing an expert. Fourth, courts must retain a robust gatekeeping role for the threshold question of whether model output is probative...", "evidence_page": null, "evidence_span": "Third, sanctions for fabricated outputs must be severe and visible, because the integrity of the entire mechanism depends on it. Recent episodes of lawyers submitting fake citations are a preview of what not doing this looks like.146 We use fabricated rather than hallucinated for a particular reason. There are subtle ways to steer models towards desired ends, and lawyer ingenuity knows no end. Litigants should be wary of risks and have the benefit of judicial penalties when opposing parties manipulate the tools in underhanded ways – much like the case of bribing an expert. Fourth, courts must retain a robust gatekeeping role for the threshold question of whether model output is probative...", "source_text_url": "https://works.battleoftheforms.com/papers/generative-gap-filling/fulltext_clean.txt", "canonical_url": "https://works.battleoftheforms.com/papers/generative-gap-filling/#claim-006", "citation": "Yonathan A. Arbel & David A. Hoffman, Generative Gap Filling, Working Paper (2026).", "topics": [], "secondary_topics": [], "human_reviewed": false, "confidence": "machine-linked", "limitations": "Machine-linked claim. Use the evidence quote and PDF before treating it as a quotation or as a complete statement of the paper's position."}
