Source-linked AI summary
Epistemic Subordination: Generative AI and the Infrastructure of Knowledge
Gilad Abiri, Emanuel V. Towfigh
TL;DR
The paper argues that generative AI structurally subordinates minority epistemologies by encoding the majority’s way of knowing as knowledge infrastructure. It analyzes this harm across legal domains and concludes that law must address the architecture where subordination is produced.
Problem
Existing legal frameworks regulate downstream decisions and applications, leaving generative AI’s architectural production of epistemic subordination insufficiently addressed.
Method
The essay traces epistemic subordination as an architectural mechanism and examines its manifestations across anti-discrimination, cultural and linguistic rights, and democratic viewpoint pluralism.
Results
The paper shows that these legal harms are three manifestations of a single architectural feature of generative AI.
Takeaways & Limitations
Law must intervene at the level where generative AI’s epistemic subordination is produced rather than only governing downstream applications.
Takeaways & Limitations
Current legal doctrines and AI-specific regulation operate downstream, governing what systems do rather than the architectural conditions producing the harm.
Abstract
from arXiv · showhide
Generative AI does not merely produce biased outputs. It encodes the majority's way of knowing as the default infrastructure of knowledge itself. We call this epistemic subordination. The training process compresses the full breadth of human expression into a single probabilistic model whose statistical baseline reflects the languages, assumptions, and cultural frameworks of the dominant culture. Minority epistemologies are not excluded but absorbed: present in the training data, yet structurally subordinated in the output. The result is not a collection of discrete biases that can be audited and corrected. It is an epistemic condition embedded in the architecture from which all outputs emerge. This unified harm cuts across three legal domains -- anti-discrimination law, cultural and linguistic rights, and democratic viewpoint pluralism -- and each fails to address it for the same structural reason: existing law regulates downstream, at the level of decisions and applications. The remedy must match the site of harm. If epistemic subordination is produced at the level of model training, then law must learn to govern at that level.
INTRODUCTION
Generative AI encodes the majority’s way of knowing as the infrastructure of knowledge, subordinating other epistemologies through model construction rather than merely producing conventional bias. Because existing legal frameworks regulate downstream outputs and applications, law must intervene in training data, alignment, and model architecture.
- INTRODUCTION: Epistemic subordination makes one epistemology the statistical default while treating every other way of knowing as a deviation.The harm is structural rather than bias in the conventional sense, although biased decisions can result.
- INTRODUCTION: Discriminatory outputs, threats to minority cultures, and narrowed democratic discourse are three manifestations of one architectural feature of generative AI.The concept unifies problems that current legal debate treats separately.
- INTRODUCTION: Internet data dominated by majority languages, assumptions, and cultural frameworks becomes the model’s default, while compression into statistical parameters prevents tracing that default to its sources.Alignment changes what models say without changing what they know.
- INTRODUCTION: Because the harm is produced during model construction, it cannot be addressed adequately through outputs, applications, or individual decisions.The essay’s central normative claim is that law must intervene at the level of construction.
- INTRODUCTION: Anti-discrimination law, cultural and linguistic rights, viewpoint pluralism, and AI-specific regulation operate downstream, governing what systems do rather than how they are built.Adequate regulation would reach training-data composition, alignment methods, and model architecture.
I. THE ARCHITECTURE OF EPISTEMIC SUBORDINATION
Generative AI turns a culturally imbalanced corpus into an opaque statistical infrastructure that encodes majority-culture epistemology as the default. This produces epistemic subordination: minority ways of knowing are structurally measured against a baseline they did not set, extending beyond discrete biased decisions.
- I. THE ARCHITECTURE OF EPISTEMIC SUBORDINATION: Generative AI is trained on the unstructured totality of online human expression, making its foundational parameters encode majority linguistic, cultural, and normative frameworks as statistical defaults.Unlike task-specific predictive systems, its training corpus is a cultural corpus shaped by digital publication demographics.
- I. THE ARCHITECTURE OF EPISTEMIC SUBORDINATION: Model opacity makes the biased baseline structurally irreversible because outputs cannot be traced to identifiable training sources or reconstructed through a causal chain.There is no variable to isolate, input to remove, or transparent decision-tree rule to inspect.
- I. THE ARCHITECTURE OF EPISTEMIC SUBORDINATION: Alignment techniques adjust model outputs but do not restructure the epistemic foundation on which the model operates.This limitation applies across reinforcement learning from human feedback, constitutional AI, system instructions, and content moderation.
- I. THE ARCHITECTURE OF EPISTEMIC SUBORDINATION: Gemini’s racially diverse Nazi soldiers exposed a statistical baseline without a framework for contextual diversity.The example illustrates how attempts to address cultural bias directly can produce historically absurd outputs.
- I. THE ARCHITECTURE OF EPISTEMIC SUBORDINATION: Epistemic subordination subordinates entire forms of knowledge and culture, implicating anti-discrimination law alongside protections for culture, language, religion, and democratic epistemic foundations.Other ways of knowing, reasoning, and communicating are measured against a majority baseline they had no part in setting.
II. THE LEGAL IMPACT OF EPISTEMIC SUBORDINATION
This section traces epistemic subordination across anti-discrimination law, language and cultural rights, and religious freedom. Although these domains differ doctrinally, each shares the purpose of protecting minorities from majority dominance, yet existing doctrine fails to reach the harm.
- II. THE LEGAL IMPACT OF EPISTEMIC SUBORDINATION: Although these legal domains operate through different doctrines and ideas, they share the purpose of protecting minorities from majority dominance.Their common general purpose links otherwise distinct legal frameworks.
- II. THE LEGAL IMPACT OF EPISTEMIC SUBORDINATION: Generative AI disrupts these protective mechanisms in different ways across the three domains.The section follows how epistemic subordination affects each domain’s mechanisms.
- II. THE LEGAL IMPACT OF EPISTEMIC SUBORDINATION: The section examines epistemic subordination across anti-discrimination law, language and cultural rights, and religious freedom.These are the three legal domains through which the paper traces the impact of epistemic subordination.
- II. THE LEGAL IMPACT OF EPISTEMIC SUBORDINATION: Existing doctrine fails to reach the harm produced by epistemic subordination.The section identifies this failure in each of the three legal domains.
A. Anti-Discrimination: The Invisible Baseline
Generative AI’s discriminatory outputs reflect model-wide baselines rather than isolated malfunctions, including dialect stereotypes, medical overassessment, unequal diagnostic recommendations, and stereotyped image associations. Anti-discrimination law struggles because its downstream tools require an identifiable practice, actor, criterion, and neutral baseline, whereas generative AI distributes harm across probabilistic models and billions of parameters.
- Discriminatory Outputs: Generative AI’s baseline associates African American English with negative stereotypes, including “dirty,” “lazy,” and “aggressive.”The descriptors are generated in response to the dialect itself.
- Discriminatory Outputs: Six to seven times the clinical baseline: medical AI systems evaluate LGBTQIA+ patients for mental health concerns at that rate.The systems also recommend less aggressive diagnostic imaging for lower-income patients.
- Discriminatory Outputs: Text-to-image generators amplify demographic stereotypes at scale by associating high-status professions with lighter skin and criminality with darker skin.These outputs indicate that discriminatory patterns are expressed through the model’s baseline rather than isolated malfunctions.
- Doctrinal Limits: Disparate-impact law can scrutinize disproportionate outcomes without discriminatory intent, but it still requires an identifiable practice that produces the disparity.Requirements such as a diploma, credit threshold, or algorithmic variable can be isolated and tested, unlike generative AI’s whole-model output.
- Doctrinal Limits: Generative AI dissolves anti-discrimination law’s presumed decision structure: outputs are probabilistic, responsibility spans multiple entities, and criteria consist of billions of statistically opaque parameters.Accordingly, intent, causation, and disparate impact lose their grip, while harm extends beyond allocative effects to inaccurate or stereotypical representations.
B. Cultural, Linguistic, and Religious Rights: The Threat from Within
Cultural, linguistic, and religious rights protect minority institutions by creating space for minority cultures to function on their own terms. Generative AI threatens this structure because cultural bias resides in model architecture, which can import majority defaults into protected institutions from within.
- B. Cultural, Linguistic, and Religious Rights: The Threat from Within: These rights share a protective mechanism: they carve out institutional space for minority cultures to operate beyond majority norms and assumptions.Examples include bilingual education, religious exemptions, and institutions maintained according to indigenous or minority traditions.
- B. Cultural, Linguistic, and Religious Rights: The Threat from Within: Generative AI reverses earlier technologies’ model of cultural orientation: bias resides in the architecture rather than selectable content.Minority communities can curate books, curricula, and online resources, but they cannot produce foundation models requiring hundreds of millions of dollars and massive data.
- B. Cultural, Linguistic, and Religious Rights: The Threat from Within: approximately seventy percent accuracy in English but only forty percent in languages such as Swahili demonstrates the multilingual disparity, which widens for smaller models.Most minority languages and cultures lack text and data at the scale available to dominant cultures.
- B. Cultural, Linguistic, and Religious Rights: The Threat from Within: AI-assisted curricula and research can import majority normative frameworks into minority institutions, including secular-Western reasoning and baselines treating minority languages as marginal.This internal restructuring of knowledge presents a challenge that cultural, linguistic, and religious protections were not designed to meet.
C. Viewpoint Pluralism and Democratic Deliberation
Epistemic subordination threatens democratic deliberation by homogenizing the epistemic field within which public discourse occurs. Existing speech and media protections cannot address this training-level harm, which requires preserving the diverse epistemologies from which pluralistic speech arises.
- Epistemic homogenization: Generative AI compresses human expression and culture into a probabilistic model whose defaults track the median culture.Minority frameworks remain in training data but are structurally subordinated in model outputs.
- Democratic deliberation: As knowledge infrastructure, generative AI will increasingly determine democratic deliberation’s baseline and constrain views expressed within public debate.Public discourse continues, but within an epistemic field already homogenized by model training.
- Limits of existing law: Speech rights and media regulation do not address the epistemic infrastructure through which generative-AI speech is produced, framed, and understood.These frameworks target speakers’ access or editorial choices by identifiable actors, whereas generative AI has no editorial gate or discrete decision to review.
- Legal response: The legal challenge is to preserve diverse epistemologies and lifeworlds, not merely protect diverse voices or ensure their public reach.Epistemic subordination collapses the assumption that diverse knowledge exists independently and only needs a channel to reach the public.
III. FROM OUTPUT TO INFRASTRUCTURE
Existing legal frameworks regulate AI mainly through downstream applications and outcomes, while epistemic subordination is embedded upstream in model-training infrastructure. Protecting epistemic pluralism therefore requires regulation of training, and emerging technical proposals show that such intervention is feasible, though legal and political frameworks remain absent.
- The regulatory gap: Epistemic subordination is an infrastructure-level condition, not an outcome, so laws regulating discriminatory decisions, minority institutions, or viewpoint diversity fail to reach its source.The harm is produced upstream during model training, while existing legal domains regulate downstream effects.
- The regulatory gap: Because training sets the epistemic baseline, preserving epistemic pluralism requires intervention in how models are built rather than only in what they produce.Relevant intervention points include training-data composition and curation, value-alignment methods, and foundation-model architecture.
- Limits of existing law: Existing training-related rules address different objectives: the EU AI Act treats dataset representativeness as product safety, while GDPR data protection can coexist with cultural homogenization.Thus, neither framework directly targets epistemic subordination.
- Technical possibilities: Emerging research proposes modular community-specific models, culturally grounded datasets, synthetic cultural data, and pluralistic alignment to preserve distinct epistemic perspectives.These approaches aim to avoid merging diverse knowledge, norms, and values into a single majority-default output.
- Technical possibilities: These proposals are unfinished, but they demonstrate that training-level intervention is technically real and developing rapidly; what remains absent is enabling and requiring legal and political support.The paper therefore identifies a gap in political will and legal framework rather than technical capacity alone.
CONCLUSION
The paper concludes that epistemic subordination exposes a common structural gap across three legal domains: law against exclusion cannot reach subordination produced through inclusion. Because the harm arises during model training, the remedy must govern data composition, alignment design, and model architecture.
- CONCLUSION: Across anti-discrimination law, cultural and linguistic rights, and viewpoint pluralism, existing protections cannot address epistemic subordination’s structural operation.Anti-discrimination law seeks a biased decision; cultural and linguistic rights protect institutional space; viewpoint pluralism promotes diverse speech, while epistemic subordination operates within and narrows the field where speech is formed.
- CONCLUSION: Epistemic subordination absorbs minority epistemologies into a majority default rather than operating through discrete acts of exclusion.This explains why legal frameworks built to protect against exclusion fail to reach the technology’s mode of subordination by inclusion.
- CONCLUSION: If epistemic subordination is produced during model training, law must govern at that level rather than only downstream applications.The proposed intervention targets the composition of data, the design of alignment, and the architecture of the model.
- CONCLUSION: Technical foundations for intervention are emerging, but the political and legal foundations needed to build them remain unfinished.Building those political and legal foundations is identified as the task ahead.