Source-linked AI summary
The Deontic Gap: Large Language Models and the Modal Language of Obligation
Daniel Hart, Sarah Allred, Joseph Abbas, Morenike Alugo
TL;DR
The paper asks whether LLMs use obligation language like contemporary humans and tests this across diverse corpora, replications, and models. It finds that LLM-generated text consistently underuses positive deontic modals, especially constructions associated with interpersonal stance.
Problem
Evidence remains limited on whether LLMs reproduce contemporary human patterns of deontic modal usage and interpersonal obligation stance.
Method
The study compares AI-generated and human text across three contemporary corpora, independent and controlled replications, historical calibration, and an eleven-model naturalistic replication.
Results
LLMs consistently underuse positive deontic modals relative to humans across primary corpora, replications, and benchmark domains.
Takeaways & Limitations
The deontic gap concentrates in interpersonal constructions such as should, have to, and had to, while modal profiles remain genre-conditional.
Takeaways & Limitations
The study does not exhaust relevant contexts, and its targeted phrase lists and deontic-versus-epistemic distinctions leave some modal uses unresolved.
Abstract
from arXiv · showhide
Modal auxiliaries such as must, should, and have to mark necessity and obligation within the contexts of speaker authority and interpersonal stance. We examine whether large language models (LLMs) reproduce contemporary human patterns of deontic modal usage. Across three primary corpora, an external benchmark, two controlled replications, and a naturalistic eleven-model replication, AI-generated text consistently underuses positive deontic modals (must, should, have to, had to) relative to contemporary humans. Historical comparison with the Google Books Ngram corpus (1920-2022), used as a heuristic calibration against the published-prose record, shows that AI modal frequencies fall within the range of formal published English, whereas contemporary human modal rates in informal digital contexts often exceed twentieth-century book baselines. Phrase-level decomposition shows that the AI-human modal gap is concentrated in constructions central to interpersonal stance (should, have to, had to), while AI matches or exceeds humans on need to in instructional and question-answering contexts but not in persuasive student writing, indicating that the modal profile is genre-conditional. The findings suggest that LLM modal usage reflects the formal written resources on which these models were trained, while underusing the modal constructions through which contemporary human writers mark immediate, interpersonal obligation.
1. Introduction
Deontic modals express obligation while also signaling differences in force, authority, and interpersonal stance. The paper’s central concern is whether LLMs reproduce these socially situated patterns rather than merely normative language in general.
- Deontic modality: Deontic modals mark duty and requirement while calibrating what is required, recommended, expected, or simply permitted.Examples include should for brushing teeth, must for returning library books, and have to for behaving in class.
- Interpersonal stance: Choosing among deontic modals represents obligation and positions speakers and listeners in social space.Have to can claim a parent’s footing, whereas must may invoke an official or institutional authority.
- Interpersonal stance: The paper’s core finding concerns differences in interpersonal stance rather than an overall poverty of normative language.This distinction between what a modal asserts and the interpersonal footing it claims organizes the paper’s analysis.
- Modal distinctions: Should commonly formulates recommendation or morally inflected advice, have to and had to locate obligation in practical constraints, and must can sound formal or institutional.The constructions differ not only in semantic strength but also in the speaker stance and interactional footing they make available.
- Register and genre: Need to frequently frames necessity as task-oriented in instructional texts, while negative forms such as can’t tend to sound conversational and immediate.Cannot is more typical of formal, institutional, or technical prose, making frequency differences relevant to corpus pragmatics.
- Model language: LLM modal usage is shaped by the register norms absorbed from training data and further adjusted by instruction-following fine-tuning.Relevant training sources include published books, academic papers, structured web content, and large-scale web scrapes.
2. Method
The study compares AI-generated and human-authored text across three contemporary genres, an external benchmark, and historical Google Books baselines. It operationalizes targeted deontic constructions, measures corpus-level rates per 10,000 words, and analyzes AI–human gaps across datasets and phrases.
- Historical calibration: Google Books Ngram data from 1920–2022 provide era-matched published-prose and published-fiction baselines for contextualizing contemporary modal rates.The general American English corpus serves as the century-long baseline, while English Fiction supplies a genre-specific comparison for fiction.
- Operationalization: Positive deontic modals comprise must, ought to, need to, has to, have to, had to, and should; rates are detected case-insensitively with whole-word boundaries and reported per 10,000 words.Negative deontic modals and impossibility terms were defined separately, with impossibility treated as a secondary exploratory comparison.
- Robustness analyses: Additional analyses tested positive-modal rate as a human/AI classifier and compared GPT-4o-to-human frequency ratios for the 300 most frequent words.The frequency-ratio analysis used paired WritingPrompts, ASAP-AES, and ELEPHANT advice-question corpora.
3. Results · 3.1. Overall Differences in the Pragmatic Marking of Obligation
Across all primary corpora and the ASAP-AES controlled replication, AI-generated text used positive deontic modals at lower rates than human-authored text. The reported AI rates were below human rates in every comparison.
- 3.1. Overall Differences in the Pragmatic Marking of Obligation: AI used positive modals at lower rates than humans across all three primary corpora and the ASAP-AES replication.The comparison covered must, should, and have to.
- 3.1. Overall Differences in the Pragmatic Marking of Obligation: 9.9 positive modals per 10,000 words was the AI rate for gsingh, compared with 13.5 for humans.This was one of four reported corpus-level comparisons.
- 3.1. Overall Differences in the Pragmatic Marking of Obligation: 21.8 positive modals per 10,000 words was the AI rate for HC3, compared with 29.2 for humans.The AI rate remained below the human rate.
- 3.1. Overall Differences in the Pragmatic Marking of Obligation: 7.6 positive modals per 10,000 words was the AI rate for WritingPrompts, compared with 25.7 for humans.This comparison showed a larger human-AI difference than gsingh or HC3.
- 3.1. Overall Differences in the Pragmatic Marking of Obligation: 25.3 positive modals per 10,000 words was the AI rate for ASAP-AES, compared with 52.9 for humans.ASAP-AES was the controlled replication sample reported in the passage.
- 3.1. Overall Differences in the Pragmatic Marking of Obligation: The reported paired values were gsingh 0.10/0.44, HC3 1.39/1.25, WritingPrompts 1.93/0.91, and ASAP-AES 2.10/0.89.The passage also reports M = 1.38/0.87.
3.2. Robustness of the Paired-Design Gaps Below the Corpus Level
Within-document analyses test whether pooled positive deontic-modal gaps are driven by a small number of documents or prompts. Across paired corpora, the evidence indicates that human and AI modal usage differs at the document and prompt levels.
- ASAP-AES: Across 149 paired ASAP-AES essays meeting the 50-word threshold, human writers averaged 65.0 positive modals per 10,000 words.The human per-document median was 34.9, with an IQR reported in the passage but truncated here.
- Paired-design robustness: The paired-corpus design enables direct within-document checks of whether aggregate modal differences are concentrated in a small number of documents or prompts.The checks cover ASAP-AES student essays and WritingPrompts Reddit creative fiction.
- WritingPrompts: In WritingPrompts, most GPT-4o stories contained no positive deontic modal, while human stories ranged widely and exceeded AI usage in most pairs.The comparison is prompt-matched, with one human-authored and one GPT-4o-generated story per prompt.
3.3. Out-of-Sample Replication Across Domains
The Pangram benchmark replicates the corpus-level finding that human-authored documents use more positive deontic modals than AI-generated documents. The difference holds across all ten domains but weakly separates individual documents.
- Out-of-Sample Replication Across Domains: 23.6 versus 14.1 positive modals per 10,000 words: human-authored Pangram documents exceeded AI-generated documents (W = 393,842, p < .001).The benchmark included 1,925 documents: 997 human and 928 AI, spanning ten domains and eight large language models.
- Out-of-Sample Replication Across Domains: Human writers exceeded AI in positive-modal mean rate across all ten Pangram domains.The cross-domain direction was consistent, but positive-modal rate separated individual human and AI documents only weakly.
3.4. Controlled Replication on Student Essays
The controlled replication tests the student-writing domain finding with a prompt-matched design. It addresses the Pangram benchmark’s limitations while sampling 150 essay prompts from the ASAP-AES corpus.
- The Pangram benchmark supports out-of-sample validation but is not a prompt-matched pragmatics corpus.
- Among better-sampled Pangram domains, student writing shows the largest human–AI positive-modal gap.The email domain has a numerically larger gap but far fewer documents.
- The controlled replication sampled 150 essay prompts from the Hewlett Foundation ASAP-AES corpus.
3.5. Historical Calibration Against Published-Prose Register
Historical calibration places AI modal rates within the range of formal edited prose, while contemporary human rates exceed historical book baselines. The comparison is heuristic rather than evidence that AI reproduces any specific decade’s language.
- Historical calibration: AI rates align more closely with the formal instructional register of published prose than contemporary human modal rates do.The narrower gap in HC3 likely reflects its informational, question-answering genre and its congruence with human expert writing.
- Historical calibration: Contemporary human writers use modals at frequencies above every historical book baseline, whereas AI output remains within the formal published-prose range.Google Books Ngram proportions were converted directly to per-10,000-word rates so AI, human, and historical frequencies shared an identical scale.
- Historical calibration: Matched years function as approximate coordinates on the historical published-prose scale, not as claims that AI encodes the language of a specific decade.The Google Books series records edited, published-text norms, which have always differed from informal human writing.
3.6. Phrase-Level Differences in Deontic Force and Stance
Phrase-level decomposition shows that AI–human differences in positive deontic modal use are concentrated in interpersonal-obligation constructions, while AI matches or exceeds humans on practical-necessity modals.
- Phrase-Level Differences in Deontic Force and Stance: The aggregate modal gap masks construction-specific differences: AI underuses have to, should, had to, and has to relative to humans.Averaged across the gsingh and HC3 datasets, humans used have to at 5.3 per 10,000 words versus 1.7 for AI.
- Phrase-Level Differences in Deontic Force and Stance: Humans use interpersonal-obligation modals more than AI, especially forms associated with recommendation, personal constraint, and narrative compulsion.The constructions are should, have to, had to, and has to.
- Phrase-Level Differences in Deontic Force and Stance: For practical-necessity modals need to and must, AI matches or exceeds human rates.This opposite pattern contrasts with the human advantage on interpersonal-obligation constructions.
3.7. Illustrative Contrasts in Modal Stance
A prompt-matched WritingPrompts contrast shows that humans and AI can express the same impossibility while conveying different emotional resonance through modal usage.
- Prompt-matched contrast: In a WritingPrompts pair about a mother fearing her daughter is becoming a “Disney princess,” both writers indicate that the situation cannot continue.The prompt involves singing, attracting animals, and receiving repeated proposals from princes.
- Prompt-matched contrast: Despite expressing the same underlying impossibility, the writers’ modal choices convey different emotional resonance.The contrast is presented as an illustrative phrase-level pattern.
- Prompt-matched contrast: The human writer stages the conversation as panicked in the example dialogue.The passage begins the human utterance with “Mary, you can’t seriously”.
3.8. Historical Change and the AI Modal Deficit
Historical change in published English explains AI–human gaps for negative modals but not for positive modals overall. Although need to and must fit the register account, most positive modals show the opposite pattern.
- Negative modals: For negative modals, historical change was positively correlated with the AI–human gap (r = .26), as predicted by the register account.Constructions that rose in books, including could not and might not, showed near-zero gaps.
- Positive modals: For positive modals, the historical-change correlation was weakly negative (r = -.31), contrary to the simple register account’s prediction.The register account predicts that constructions rising in books should be used by AI at near-human rates.
- Positive modals: Need to fits the register prediction after a 620% rise in Google Books since 1960 and has the largest AI-favoring gap among positive modals.The rise was independently documented in reference corpora by Leech et al. (2009).
- Positive modals: Must fits the register prediction weakly, with a roughly flat historical trajectory and a modest AI advantage.The passage contrasts must with the other five positive modals, which do not fit the simple register account.
3.9. Generality Across Model Families
The deontic gap generalizes across independently developed model families: Claude, like GPT-4o, uses positive deontic modals far less often than contemporary human writers under identical conditions.
- 3.9. Generality Across Model Families: Claude used positive deontic modals at 9.7 per 10,000 words, versus 25.7 for contemporary humans and 7.6 for GPT-4o.The comparison comes from 150 WritingPrompts stories generated from the same prompts under identical conditions, with only the model differing between AI columns.
- 3.9. Generality Across Model Families: Both independently developed model families fall far below the contemporary human rate in the matched WritingPrompts comparison.The result supports generality across model families rather than a pipeline-specific effect.
- 3.9. Generality Across Model Families: The gap reflects a specific and repeatable feature of how models mark obligation, rather than a broad failure to render common vocabulary.This conclusion follows from the model-family comparison described in the supplied passages.
4. Discussion
Across datasets, benchmarks, controlled replications, and eleven-model comparisons, LLMs use deontic obligation language differently from contemporary humans in a patterned way. The gap concentrates in interpersonal obligation constructions and is consistent with formal-register exposure, procedural training data, and measured advisory preferences.
- Core finding: Across three corpora, an independent benchmark, controlled replications, and eleven models, AI consistently underuses positive deontic modals relative to contemporary humans.The study reports this pattern across multiple datasets, replication designs, and model families.
- Register account: WritingPrompts humans use positive modals at 25.7 per 10,000 words, versus 7.6 in GPT-4o responses to the same prompts.This comparison illustrates the register asymmetry between contemporary informal fiction writing and AI-generated text.
- Implications and limitations: The authors argue that AI has absorbed formal written norms from training data, while contemporary human informal language has continued to informalize beyond those modal-density patterns.They frame the difference as register fixation rather than random noise, while noting that the available corpora do not exhaust all contexts.
- Core finding: The largest AI–human gaps occur for should, have to, and had to, which express recommendation, compulsion, and retrospectively obligatory behavior.These constructions convey especially strong and personally immediate senses of obligation.
- Implications and limitations: The established AI–human gap motivates research on whether AI-generated guidance affects readers’ interpretations and becomes part of the linguistic environment shaping future speakers.The paper connects this agenda to evidence that LLM-favored lexical patterns are entering human speech.
Supplementary Material
The supplementary material documents the corpus-rate methodology and benchmarks, and charts historical trajectories and per-document modal-rate distributions for human and AI texts.
- ELEPHANT Advice Corpus: Supplementary Table S1 reports positive deontic-modal rates and length-adjusted rate ratios for human writers and eleven language models in the ELEPHANT advice corpus.Rates count occurrences per 10,000 words, pooled over responses of at least 50 words.
- Rate-Ratio Method: The rate ratio estimates each model’s multiplicative difference from the human rate using a Poisson generalized linear mixed model with word-count offsets and prompt-level random intercepts.The adjustment accounts for the greater length of machine-generated responses.
- Historical Trajectories: Figure S1 indexes each positive deontic construction in Google Books from 1920–2022 to its own historical mean, enabling comparison across constructions with different base rates.Need to rises steadily; have to, had to, and has to rise mainly in recent decades, while should and must remain essentially flat.
- Per-Document Distributions: Figure S2 compares per-document positive deontic-modal rates for human and GPT-4o texts in prompt-matched WritingPrompts fiction and ASAP-AES student essays.The distributions include documents of at least 50 words, with dashed writer-specific median lines and annotations for each panel’s median and zero-rate percentage.