# Propositions from The Generative Reasonable Person

**Citation:** Yonathan A. Arbel, The Generative Reasonable Person, BYU Law Review (forthcoming 2027), arXiv:2508.02766v2 (2026)

**Source:** [arXiv v2 manuscript PDF](https://arxiv.org/pdf/2508.02766)

**Review status:** 12 model-drafted, source-checked; 0 human-reviewed. Page references use the printed pagination and, separately, the 1-based PDF page number.

## 1. The generative reasonable person supplies an empirical reference point for legal judgments that invoke ordinary reasonableness

**Location:** Introduction, printed pp. 2-7 (PDF pp. 2-7)

Professor Yonathan A. Arbel claims, in “The Generative Reasonable Person” on pages 2–7, that law has long invoked the views of reasonable consumers, jurors, and ordinary people without a cheap, scalable way to measure those views. He introduces the generative reasonable person as an LLM-based empirical reference point, implemented through Silicon Randomized Controlled Trials, that can test elite intuition against simulated lay judgments. This is significant because it reframes many judicial statements about what no reasonable person could believe as testable empirical bets rather than self-validating common sense. It connects to jury studies, consumer surveys, ordinary-meaning research, access to justice, and the use of dictionaries as aids that inform rather than decide legal judgment.

**Evidence anchor:** The introduction identifies the absence of an affordable empirical baseline, defines the new method, previews three replications totaling nearly 10,000 simulated judgments, and describes judicial, regulatory, and litigant uses while reserving final judgment to humans.

**Boundary:** Arbel presents the tool as an empirical starting point and a cautious improvement over current methods, not as an arbiter, a complete substitute for human subjects, or proof that every model can reproduce every lay judgment.

**Connections:** reasonable-person doctrine; experimental jurisprudence; ordinary meaning; access to justice; judicial intuition

**Record:** `ssrn-5377475-p01` · `machine-drafted-source-checked`

## 2. Lay judgments remain relevant to reasonableness even when they do not control the normative legal standard

**Location:** Part I, Folk Opinions and the Law, printed pp. 8-10 (PDF pp. 8-10)

Professor Yonathan A. Arbel claims, in “The Generative Reasonable Person” on pages 8–10, that law is simultaneously a professional system and a social institution that must remain attentive to the language, experience, and judgments of the governed. Lay views matter reflectively because they illuminate legal concepts, pragmatically because law must guide conduct, democratically because comprehensibility supports legitimacy and participation, and epistemically because dispersed lived experience contains information elites may lack. This is significant because it avoids the false choice between treating public opinion as dispositive and treating it as irrelevant. It connects to folk jurisprudence, the plain-language movement, ordinary-meaning interpretation, civil juries, Hayekian dispersed knowledge, and hybrid theories in which descriptive facts inform but do not determine normative conclusions.

**Evidence anchor:** Part I traces the law’s recurring reliance on common language and community standards, offers reflective, effectiveness, legitimacy, political, and informational reasons to consult lay views, and ends by asking how the state can make those views legible.

**Boundary:** The claim is not that majorities define legal rightness; Arbel expressly maintains that the descriptive baseline can inform a normative decision without controlling it.

**Connections:** folk jurisprudence; democratic legitimacy; plain language; civil jury; dispersed knowledge; descriptive and normative reasonableness

**Record:** `ssrn-5377475-p02` · `machine-drafted-source-checked`

## 3. LLM architecture makes simulated lay judgment plausible while creating predictable majoritarian, granular, and temporal limits

**Location:** Part II, Generative People in Theory and the Social Sciences, printed pp. 11-17 (PDF pp. 11-17)

Professor Yonathan A. Arbel claims, in “The Generative Reasonable Person” on pages 11–17, that attention, roleplaying, generalization, and a statistical tendency toward common patterns make modern LLMs plausible instruments for approximating ordinary judgments. The same majoritarian tendency that helps a model recover widespread social schemas can reproduce entrenched bias, flatten minority perspectives, simulate particular people poorly, and become stale as norms change. This is significant because the article derives both the promise and the principal risks of the method from the same underlying machinery rather than treating bias as an unrelated implementation defect. It connects to transformer attention, silicon sampling, persona research, generalization, the feminist critique of the reasonable man, group stereotyping, and temporal value drift.

**Evidence anchor:** Part II links four model capabilities to the proposed use, surveys evidence of human-like judgments and roleplay, and identifies bias amplification, granularity, stereotyping, speaking-for concerns, and aging training data as structural cautions.

**Boundary:** Architectural plausibility and prior social-science studies do not establish legal validity; roleplaying quality depends on context, demographic prompts can stereotype, and apparent demographic alignment may be a prompting artifact.

**Connections:** transformer models; roleplaying; silicon sampling; majoritarian bias; minority perspectives; value drift

**Record:** `ssrn-5377475-p03` · `machine-drafted-source-checked`

## 4. Silicon Randomized Controlled Trials use stateless sessions and differential measurement to test latent model sensitivity rather than doctrinal recall

**Location:** Part III.1, The Core Methodology: S-RCT, printed pp. 18-20 (PDF pp. 18-20)

Professor Yonathan A. Arbel claims, in “The Generative Reasonable Person” on pages 18–20, that a valid test of simulated reasonableness must separate internalized lay patterns from memorized cases and responses tailored to the researcher’s expectations. His Silicon Randomized Controlled Trial method randomly assigns conditions across fresh, stateless model sessions, measures changes between conditions instead of equating raw human and model scores, adds personas as an experimentally testable treatment, and checks results across models. This is significant because it adapts causal-inference logic to an artificial subject while directly addressing contamination, sycophancy, cross-condition harmonization, and scale calibration. It connects to randomized controlled trials, between-subjects design, counterfactual evaluation, model ablation, persona prompting, and robustness across proprietary and open architectures.

**Evidence anchor:** The methodology section identifies recall and sycophancy as separate validity threats, then specifies three mechanisms—independent sessions, differential outcomes, and persona assignment—plus cross-model tests for robustness.

**Boundary:** Fresh sessions reduce but cannot eliminate contamination or demand effects, persona effects must be validated rather than assumed, and directional concordance does not by itself prove quantitative calibration or human equivalence.

**Connections:** randomized controlled trials; causal inference; stateless API sessions; differential measurement; persona ablation; cross-model validation

**Record:** `ssrn-5377475-p04` · `machine-drafted-source-checked`

## 5. The negligence replication recovered the lay priority of social conformity over cost-benefit analysis but overstated effect magnitudes

**Location:** Part III.2.1, The Empirical Reasonable Person, printed pp. 20-27 (PDF pp. 20-27)

Professor Yonathan A. Arbel claims, in “The Generative Reasonable Person” on pages 20–27, that models reproduce the counter-doctrinal hierarchy found in Christopher Jaeger’s negligence experiment: people react more strongly to whether a precaution is common than to whether it is economically justified. The S-RCT retained 5,529 of 5,544 planned responses; pooled persona-model judgments moved about 9.71 points with commonness and 4.41 points with cost, compared with human effects of roughly 4.98 and 1.11 points. This is significant because the shared ordering suggests that models captured a lay social schema rather than merely reciting the Hand Formula or doctrine minimizing custom. It connects to negligence, customary practice, economic analysis of torts, social-norm theory, experimental replication, and the need to distinguish qualitative structure from quantitative calibration.

**Evidence anchor:** The section reproduces a two-by-two commonness-and-cost design, reports high response retention and pooled effects, compares them with Jaeger’s human results, and analyzes model-level and persona-ablation variation.

**Boundary:** The models were substantially more sensitive than humans to both manipulations, the human economic estimate was not statistically significant while the silicon estimate was, one model did not preserve the hierarchy, and persona benefits varied by model and dimension.

**Connections:** negligence; Hand Formula; custom and social norms; human-subject replication; effect-size calibration

**Record:** `ssrn-5377475-p05` · `machine-drafted-source-checked`

## 6. Models replicated the lay paradox that an essential lie undermines consent more than a material lie that matters more to the victim

**Location:** Part III.2.2, Generative Commonsense Consent, printed pp. 27-32 (PDF pp. 27-32)

Professor Yonathan A. Arbel claims, in “The Generative Reasonable Person” on pages 27–32, that LLMs reproduce Roseanna Sommers’s counterintuitive structure of consent under deception. Across 3,232 judgments from 202 synthetic personas, the pooled models treated a lie about reward points as more important to the buyer yet still perceived more consent than when the seller lied about the identity of the product; seven of eight models reproduced each directional effect. This is significant because canonical doctrine emphasizes materiality, whereas the repeated model pattern suggests that ordinary people separately privilege authenticity about the transaction’s essence. It connects to consent theory, fraudulent misrepresentation, transaction identity, commonsense moral schemas, model safety training, and domain-specific calibration.

**Evidence anchor:** The experiment randomizes personas between essential- and material-lie scenarios, measures consent and perceived importance, reports pooled and model-level directional replication, and documents compression, ceiling effects, and persona ablation results.

**Boundary:** The model consent shift was compressed and importance ratings often reached the scale ceiling; failures differed by model and prompt, persona effects were heterogeneous, and the closest-calibrated model differed from the negligence study.

**Connections:** commonsense consent; fraud and misrepresentation; material versus essential deception; transactional authenticity; silicon replication

**Record:** `ssrn-5377475-p06` · `machine-drafted-source-checked`

## 7. In the hidden-fee study, models reproduced lay contract formalism and usually fell nearer lay than elite legal baselines

**Location:** Part III.2.3, Generative Language Sense, Fairness Sense, and Legal Sense, printed pp. 33-37 (PDF pp. 33-37)

Professor Yonathan A. Arbel claims, in “The Generative Reasonable Person” on pages 33–37, that models reproduce the lay tendency to separate fairness, consent, and anticipated legal enforcement in a deceptive hidden-fee contract. All eight models ranked the fee lowest on fairness, higher on consent, and highest on likely enforceability; twenty-three of twenty-four model-by-question means fell within one standard deviation of lay baselines, and five of eight models were nearer the lay three-dimensional profile than the legal-professional profile. This is significant because it tests not only whether models move like humans but whose absolute evaluative voice they most resemble. It connects to contract formalism, fine-print fraud, consumer consent, calibration against competing populations, and the concern that legal AI might disguise elite professional judgment as public sentiment.

**Evidence anchor:** Using an instrument previously administered to lay and legally trained respondents, the study compares model means for fairness, consent, and enforcement, tests their hierarchy and Euclidean distance to each human benchmark, and runs persona ablations.

**Boundary:** The five-of-eight tendency toward the lay cluster was not statistically significant, dimension-level results were mixed, fairness scores leaned toward lawyers, and persona movement toward lay baselines was modest and not uniform across models.

**Connections:** lay contract formalism; fine-print fraud; consumer consent; lay-lawyer comparison; persona calibration

**Record:** `ssrn-5377475-p07` · `machine-drafted-source-checked`

## 8. Models are better supported as maps of what tends to matter in lay judgment than as precision forecasters of how much it matters

**Location:** Part IV.1, Interpretation of the Findings, printed pp. 38-42 (PDF pp. 38-42)

Professor Yonathan A. Arbel claims, in “The Generative Reasonable Person” on pages 38–42, that the three replications reveal an internal geometry of lay reasonableness: structural relationships among social conformity and cost, essential and material deception, and fairness, consent, and enforceability recur across domains and architectures. At the same time, alignment training and other model features appear to amplify some effects and compress or cap others. This is significant because it defines a narrower but more defensible use—comparing directions, rankings, and sensitivity to factors—than treating model ratings as calibrated population estimates. It connects to construct validity, qualitative versus quantitative replication, reinforcement learning from human feedback, sensitivity analysis, and the evidentiary difference between identifying a relevant factor and estimating its precise weight.

**Evidence anchor:** Part IV synthesizes the replicated hierarchies, emphasizes consistency across domains and architectures, catalogs magnitude distortions, and contrasts supported comparative uses with unsupported numerical forecasting.

**Boundary:** Only three published human studies and selected dimensions were replicated; leakage, demand effects, multiple testing, random chance, and flaws or limited generalizability in the underlying human studies cannot be fully excluded.

**Connections:** construct validity; directional replication; quantitative calibration; alignment training; sensitivity analysis

**Record:** `ssrn-5377475-p08` · `machine-drafted-source-checked`

## 9. The proper role of simulated lay judgment depends on whether a legal standard is descriptive, normative, or hybrid

**Location:** Part IV.2, Domains of Application: Between Promise and Prudence, printed pp. 42-43 (PDF pp. 42-43)

Professor Yonathan A. Arbel claims, in “The Generative Reasonable Person” on pages 42–43, that legal reasonableness inquiries should be sorted into explicitly descriptive, explicitly normative, and hybrid domains before model evidence is assigned a role. Simulated public understanding bears most directly on descriptive tests such as reasonable-consumer deception, should function only as a transparency check for normative tests such as constitutional balancing, and can supply an empirical predicate without resolving the prescriptive conclusion in hybrid fields such as negligence, consent, and contract interpretation. This is significant because it prevents the availability of cheap empirical output from silently converting moral or constitutional questions into opinion polls. It connects to doctrinal fit, law-fact boundaries, consumer protection, the Hand Formula, constitutional reasonableness, consent, and objective contract interpretation.

**Evidence anchor:** The applications section defines three zones, gives doctrinal examples of each, and specifies that silicon evidence answers descriptive predicates while human decisionmakers retain prescriptive authority.

**Boundary:** The categories can overlap and require legal interpretation; classifying a standard does not solve calibration, bias, admissibility, or the ultimate normative question.

**Connections:** descriptive legal standards; normative legal standards; hybrid standards; consumer deception; constitutional balancing; law-fact distinction

**Record:** `ssrn-5377475-p09` · `machine-drafted-source-checked`

## 10. Generative reasonable people can serve as low-cost pretests and empirical guardrails for regulators, courts, litigants, and firms

**Location:** Part IV.2.1–4, Institutional Applications, printed pp. 43-47 (PDF pp. 43-47)

Professor Yonathan A. Arbel claims, in “The Generative Reasonable Person” on pages 43–47, that disciplined silicon studies can cheaply pretest public understanding for rulemaking, challenge judges’ assumptions about consumers, give under-resourced litigants a rough analogue to jury consulting, and help firms screen contracts, advertising, and compliance choices before harm or litigation. The common institutional design is tiered: use models to identify likely trouble and decide where expensive surveys, focus groups, discovery, or direct consultation are most valuable. This is significant because the relevant comparison is often not a perfect human study but no consultation at all, outdated surveys, elite intuition, or feedback distorted by money and mobilization. It connects to FTC deception policy, adversarial testing under procedures analogous to court-appointed expertise, litigation equality, preventive compliance, and staged allocation of empirical-research resources.

**Evidence anchor:** The section works through regulatory labeling, consumer cases, resource-constrained litigation, privacy and marketing compliance, and contract drafting, consistently positioning model studies as preliminary or supplementary evidence.

**Boundary:** Reliability falls with demographic granularity; high-stakes or minority-sensitive matters strengthen the case for real consultation; inexperienced users may overtrust commercial jury-prediction products; and courts must expose model and prompt choices to adversarial challenge.

**Connections:** regulatory pretesting; consumer deception; judicial discretion; access to justice; compliance design; tiered empirical research

**Record:** `ssrn-5377475-p10` · `machine-drafted-source-checked`

## 11. An accessible empirical baseline changes reasonable-person theory by forcing normative departures from public understanding into the open

**Location:** Part IV.2.5, Legal Debates Between the Descriptive and Normative Person, printed pp. 47-48 (PDF pp. 47-48)

Professor Yonathan A. Arbel claims, in “The Generative Reasonable Person” on pages 47–48, that the longstanding debate over whether the reasonable person is descriptive, normative, or hybrid has been shaped partly by the practical scarcity of reliable information about ordinary judgment. Generative reasonable people can loosen that constraint without making public opinion authoritative: descriptivists gain a measurable baseline, while normativists gain a way to test whether proposed rules are communicable and to identify when doctrine deliberately departs from public understanding. This is significant because courts could no longer present a contested policy choice as though it were simply a report about what everyone naturally thinks. It connects to legal realism, democratic accountability, administrability, expressive clarity, second-best institutional theory, and the distinction between candid normative justification and empirical rhetoric.

**Evidence anchor:** The theoretical discussion argues that empirical scarcity has organized prior positions, explains why both descriptive and normative theories benefit from a usable baseline, and focuses on making departures explicit rather than forbidding them.

**Boundary:** A simulated baseline remains contestable and does not resolve which departures are justified; normative commitments to equality, constitutional rights, efficiency, or minority protection may properly outweigh majority judgment.

**Connections:** descriptive-normative debate; legal realism; democratic accountability; administrability; normative transparency

**Record:** `ssrn-5377475-p11` · `machine-drafted-source-checked`

## 12. Legal deployment requires human authority, transparent methods, bias audits, real-community validation, triangulation, and temporal maintenance

**Location:** Part IV.2.6 and Conclusion, Principles and Best Practices, printed pp. 48-51 (PDF pp. 48-51)

Professor Yonathan A. Arbel claims, in “The Generative Reasonable Person” on pages 48–51, that generative reasonable people should augment rather than supplant human judgment and must be governed as fallible empirical instruments. He calls for disclosure of models, prompts, and personas; adversarial comparison; calibrated confidence; testing across protected and intersectional groups; continuing engagement with real minority communities; triangulation with surveys or focus groups in high-stakes settings; and attention to knowledge cutoffs and changing norms. This is significant because a model’s majoritarian reach cannot provide democratic legitimacy if its operation hides excluded voices, stale values, or false numerical precision. It connects to evidence governance, disparate-impact auditing, lived experience, Bayesian use of uncertain evidence, reproducibility, dynamic representation, and the article’s closing claim that technology can make ordinary people more legible without outsourcing legal judgment.

**Evidence anchor:** The final section lists adjunct-not-arbiter, transparency, calibration, bias, mimesis, triangulation, and dynamic-representation principles, then concludes that the contribution is scalable access to an empirical predicate rather than algorithmic adjudication.

**Boundary:** Arbel describes a preliminary roadmap rather than a definitive protocol; models cannot reproduce the phenomenological richness of lived experience, persona simulation degrades with intersectional complexity, and retraining cannot itself decide which values law should preserve.

**Connections:** human-in-the-loop governance; methodological transparency; bias auditing; community validation; triangulation; temporal drift

**Record:** `ssrn-5377475-p12` · `machine-drafted-source-checked`
