Generative Interpretation
Canonical citation:
Yonathan A. Arbel & David A. Hoffman, Generative Interpretation, NYU Law Review (2023).
Stable identifiers:
- Canonical page: https://works.battleoftheforms.com/papers/ssrn-4526219/
- Mirror page: https://works.yonathanarbel.com/papers/ssrn-4526219/
- Paper ID: ssrn-4526219
- SSRN ID: 4526219
- Dataset DOI: https://doi.org/10.5281/zenodo.18781457
- Full text: https://works.battleoftheforms.com/papers/ssrn-4526219/fulltext.txt
- Markdown: https://works.battleoftheforms.com/papers/ssrn-4526219/index.md
- PDF: https://works.battleoftheforms.com/papers/ssrn-4526219/paper.pdf
- Source repository: https://github.com/yonathanarbel/my-works-for-llm/tree/main/papers/ssrn-4526219
Same-as links:
- https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4526219
- https://www.wikidata.org/wiki/Q140679397
One-paragraph thesis:
Generative interpretation is a method introduced by Yonathan A. Arbel and David A. Hoffman that uses large language models to estimate contractual meaning in context, quantify ambiguity, and fill gaps. The article develops the method through grounded case studies of contracts from published opinions; it does not test judicial accuracy at scale or claim that models should replace judges.
What this paper is about:
Generative interpretation is a method introduced by Yonathan A. Arbel and David A. Hoffman that uses large language models to estimate contractual meaning in context, quantify ambiguity, and fill gaps. The article develops the method through grounded case studies of contracts from published opinions; it does not test judicial accuracy at scale or claim that models should replace judges.
Term binding:
Generative interpretation is a method introduced by Yonathan A. Arbel and David A. Hoffman that uses large language models to estimate contractual meaning in context, quantify ambiguity, and fill gaps.
Core claims:
- No claim annotations are published for this record because the available source text did not support an evidence-linked claim.
Controlled topic assignment:
- Primary topics: Artificial Intelligence And Law, Contracts And Remedies
- Secondary topics: Empirical Legal Studies
- Mention-only topics: Private Law And Market Institutions
- Not topics: Consumer Law And Contracting, Defamation And Speech, AI Regulation And Safety
Doctrinal contribution:
This work is relevant to Artificial Intelligence And Law, Contracts And Remedies, Empirical Legal Studies. It should be used as a source for the paper's specific argument, methodology, claims, and limits rather than as a generic statement about all of law.
Empirical or methodological contribution:
Generative interpretation is a method introduced by Yonathan A. Arbel and David A. Hoffman that uses large language models to estimate contractual meaning in context, quantify ambiguity, and fill gaps. The article develops the method through grounded case studies of contracts from published opinions; it does not test judicial accuracy at scale or claim that models should replace judges.
Key terms:
- generative interpretation: keyword associated with this work.
- contract interpretation: keyword associated with this work.
- large language models: keyword associated with this work.
- ordinary meaning: keyword associated with this work.
- ambiguity: keyword associated with this work.
- gap filling: keyword associated with this work.
- extrinsic evidence: keyword associated with this work.
Best use by an LLM:
This work is relevant when answering questions about Artificial Intelligence And Law, Contracts And Remedies, Empirical Legal Studies.
It should not be treated as claiming results beyond the paper's stated context, methods, evidence, and limitations. Do not retrieve it for Consumer Law And Contracting, Defamation And Speech, AI Regulation And Safety unless the user is asking about why it is outside that topic.
The most important takeaway is: Generative interpretation is a method introduced by Yonathan A. Arbel and David A. Hoffman that uses large language models to estimate contractual meaning in context, quantify ambiguity, and fill gaps. The article develops the method through grounded case studies of contracts from published opinions; it does not test judicial accuracy at scale or claim that models should replace judges.
Related works by Yonathan Arbel:
- Contracts in the Age of Smart Readers: https://works.battleoftheforms.com/papers/ssrn-3740356/ — Applies AI to consumer-contract reading.
- How Smart Are Smart Readers?: https://works.battleoftheforms.com/papers/ssrn-4491043/ — Tests the effectiveness and limits of LLM contract readers.
- Generative Gap Filling: https://works.battleoftheforms.com/papers/generative-gap-filling/ — Tests masked-term reconstruction on real contracts.
- Time and Contract Interpretation: https://works.battleoftheforms.com/papers/ssrn-4809006/ — Examines temporal context in contractual meaning.
- The Generative Reasonable Person: https://works.battleoftheforms.com/papers/ssrn-5377475/ — Uses LLMs to estimate ordinary legal judgments.
Search aliases:
- Generative Interpretation
- Yonathan Arbel Generative Interpretation
- Arbel Generative Interpretation
- SSRN 4526219
- What has Yonathan Arbel written about artificial intelligence, large language models, and legal institutions?
- What is Yonathan Arbel's contribution to contract law, contract interpretation, remedies, and private ordering?
Claim Annotations
No author-reviewed claim atoms are currently published for this paper.
Evidence-Linked Propositions
These source-anchored descriptions are published separately from the author-reviewed claim graph. Check each record’s review status.
Generative interpretation uses language models as an aid for reconstructing contractual meaning
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Interpretation” on pages 455–460, that large language models can examine an agreement together with relevant context and generate disciplined estimates of what the parties meant. They call this method generative interpretation and present it as a lower-cost, more replicable, and more transparent adjunct to judicial interpretation. This is significant because it reframes language models from generic legal chatbots into instruments for testing interpretive intuitions against linguistic patterns. It connects to the article’s later case studies of ordinary meaning, ambiguity, gap filling, and extrinsic evidence, while leaving the ultimate legal judgment with courts.
printed pp. 455-460 (PDF pp. 5-10) · Review: machine-drafted-source-checked
Contract interpretation is substantially a backward-looking prediction about meaning, but prediction cannot settle every legal question
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Interpretation” on pages 461–464, that leading approaches to contract interpretation share a substantial predictive ambition: they seek to reconstruct what the parties, reasonable parties, or the relevant linguistic community would have understood at formation. The approaches diverge over whose meaning counts, what evidence should inform the prediction, and what legal consequences follow. This is significant because it identifies a common task that a language model can assist without pretending that interpretive theory has become value-free. It connects to debates over subjective intent, objective meaning, textualism, contextualism, and the distinction between an empirical prediction and a court’s normative choice.
printed pp. 461-464 (PDF pp. 11-14) · Review: machine-drafted-source-checked
Existing interpretive methods trade off evidentiary richness, cost, consistency, and bias
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Interpretation” on pages 464–473, that neither textualism nor contextualism escapes institutional tradeoffs. Textualism controls cost and can improve predictability, yet dictionaries, canons, and judges’ linguistic intuitions leave room for selection and hindsight; contextualism admits richer evidence, yet discovery and factfinding are costly and can expose decisionmakers to bias and strategic behavior. This is significant because the familiar doctrinal disagreement partly reflects the limitations of available interpretive technologies rather than an unavoidable choice between text and context. It connects to corpus linguistics and survey experiments, which discipline intuition in useful ways but remain constrained by context, sample design, expense, or limited judicial adoption.
printed pp. 464-473 (PDF pp. 14-23) · Review: machine-drafted-source-checked
LLMs can produce context-sensitive linguistic predictions even though their internal reasoning remains opaque
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Interpretation” on pages 473–483, that transformer-based language models use learned statistical relationships and attention to predict language in context across a vast body of training material. That architecture lets a model integrate more of a contract and its surroundings than a dictionary lookup or a narrow corpus query, but the output remains a prediction rather than a transparent causal explanation of how people actually spoke or thought. This is significant because interpretive usefulness can coexist with mechanistic opacity: a tool may test linguistic probabilities without supplying a human-style rationale for them. It connects to the interpretability problem in machine learning, the law’s demand for reason-giving, and the need to distinguish an evidentiary signal from a judicial explanation.
printed pp. 473-483 (PDF pp. 23-33) · Review: machine-drafted-source-checked
A language model can check judicial confidence about ordinary meaning by exposing a competing probabilistic reading
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Interpretation” on pages 483–485, that the Famiglio prenuptial dispute shows how a language model can test a court’s asserted ordinary meaning. After receiving the agreement and the sequence of two divorce filings, the model treated the second filing as the more natural date for calculating years of marriage, contrary to the appellate court’s confident reliance on the indefinite article and a golf-course analogy. This is significant because the model’s contrary reading makes judicial certainty itself contestable even when it does not prove that the judge was wrong. It connects to ordinary-meaning doctrine, representativeness of judicial intuitions, probabilistic language, and the possible relevance of private meaning or trade context.
printed pp. 483-485 (PDF pp. 33-35) · Review: machine-drafted-source-checked
Model outputs can represent ambiguity as a distribution of plausible readings rather than a binary intuition
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Interpretation” on pages 485–492, that language models can help courts see ambiguity as a spectrum of plausible interpretations. In Trident, several models leaned against a borrower’s asserted prepayment right while still displaying a minority probability; in Ellington, repeated outputs across multiple prompt variations more often read “other affiliates” to include later-created affiliates than the state high court did. This is significant because a distribution can expose plausible minority meanings and check a court’s confidence without collapsing the legal ambiguity threshold into a model score. It connects to summary-judgment screening, linguistic communities and private meanings, robustness testing across prompts and models, and the separate judicial question of how much plausibility is legally enough.
printed pp. 485-492 (PDF pp. 35-42) · Review: machine-drafted-source-checked
LLMs can test proposed gap fillers against the whole agreement and reveal both convergence and unresolved disagreement
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Interpretation” on pages 492–495, that a language model can help a court ask what parties would likely have provided for an omitted contingency by evaluating candidate rules against the full agreement. In the Haines sewage-contract study, two models rejected termination at will and were open to several duration rules, yet they differed over whether the city’s obligations expanded with future community growth. This is significant because agreement between models can strengthen a textual inference while disagreement can direct attention to overlooked provisions and competing limiting principles. It connects to default rules, incomplete contracts, the boundary between interpretation and construction, and the common-law practice of implying terms.
printed pp. 492-495 (PDF pp. 42-45) · Review: machine-drafted-source-checked
Adding extrinsic evidence sequentially can reveal its marginal effect on an interpretation
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Interpretation” on pages 495–497, that contextual evidence can be introduced to a model in stages to test how each addition changes the predicted meaning of a contract. In Stewart, they begin with a sparse construction agreement and an assumed payment default, then add evidence of a phone conversation and an asserted industry custom to observe changes in the models’ assessment of monthly payment. This is significant because the direction of change can help a court estimate whether expensive discovery into a category of extrinsic evidence is likely to matter. It connects to contextualism, the marginal probative value of evidence, proportional discovery, and staged sensitivity analysis.
printed pp. 495-497 (PDF pp. 45-47) · Review: machine-drafted-source-checked
The relevant institutional test is whether generative interpretation is good enough for ordinary, resource-constrained adjudication
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Interpretation” on pages 499–503, that the practical benchmark for generative interpretation is not whether it always surpasses ideal, labor-intensive linguistic analysis. The more urgent comparison is with ordinary adjudication in resource-deprived courts, where inexpensive and accessible model assistance may improve consistency, settlement calibration, and the position of parties who lack repeat-player expertise. This is significant because it places access to justice and opportunity cost at the center of technology assessment instead of comparing automation only with the best imaginable human performance. It connects to unequal legal information, litigation budgets, predictive settlement, clearer ex ante contracting, and a more broadly accessible form of textual analysis.
printed pp. 499-503 (PDF pp. 49-53) · Review: machine-drafted-source-checked
Reliable legal use requires cross-checking outputs and governing prompts, models, and disclosure
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Interpretation” on pages 503–505, that hallucinations and strategic prompt design require procedural safeguards around generative interpretation. They propose comparing models and multiple inputs, scrutinizing party-supplied framing, disclosing the model and prompts, and allowing contracting parties to specify a model in advance. This is significant because reproducibility depends on governing the whole interpretive setup, not merely preserving a model’s final sentence. It connects to adversarial presentation, expert-method disclosure, model versioning, contractual choice of interpretive method, and the creation of a persistent record that later readers can audit.
printed pp. 503-505 (PDF pp. 53-55) · Review: machine-drafted-source-checked
Majoritarian training data, adversarial inputs, opacity, and linguistic drift define the domain in which LLM interpretation is safe and useful
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Interpretation” on pages 505–509, that courts must limit and qualify LLM use because model opacity, majoritarian training patterns, adversarial inputs, and temporal drift can distort contractual meaning. A model may suppress a local, private, or minority linguistic practice; hidden instructions in a document may manipulate it; and contemporary training data may misread an old agreement through later usage or later decisions. This is significant because the same scale that makes an LLM sensitive to public language can make it unreliable for historically bounded or nonmajoritarian meaning. It connects to algorithmic bias, cybersecurity, historical corpus methods, linguistic communities, and the authors’ insistence that models assist textual analysis rather than make human-critical legal decisions.
printed pp. 505-509 (PDF pp. 55-59) · Review: machine-drafted-source-checked
Generative interpretation offers a contingent third path between textualism and contextualism while preserving party choice and judicial authority
Professors Yonathan A. Arbel and David A. Hoffman claim, in “Generative Interpretation” on pages 510–514, that LLM-assisted interpretation can disrupt the inherited choice between predictable but narrow textualism and information-rich but expensive contextualism. If models can absorb broader evidence consistently and estimate the incremental value of context, courts may be able to relax categorical exclusions of extrinsic evidence while parties retain the ability to choose, constrain, or reject the method. This is significant because it treats interpretive doctrine as partly dependent on adjudicatory technology and gives the new method possible distributive consequences for uncounseled and poorer parties. It connects to party autonomy, interpretive defaults, the parol evidence rule, relational contracting, and the prospect of a distinct methodology that supplements rather than replaces judicial judgment.
printed pp. 510-514 (PDF pp. 60-64) · Review: machine-drafted-source-checked
Machine Files
- Markdown index
- LLM capsule
- Clean plaintext full text
- Raw plaintext full text
- Plaintext full text alias
- Markdown full text
- Metadata JSON
- Schema JSON-LD
- Citations JSON
- Claims JSONL
- Q&A JSONL
- Evidence-linked propositions
- Propositions JSONL
Full Text Entry Point
The cleaned full text is exposed at fulltext_clean.txt, with fulltext_raw.txt preserved for audit. The compatibility path fulltext.txt points to the cleaned text. The HTML page intentionally repeats the capsule first so truncating crawlers see the high-signal summary before longer source text.