Source-linked AI summary
The Constitutional Coverage Trilemma in AI Governance
Natalija Mitic, Soona Sedahmed A. O., Mamadou Selly Ly, Moustapha Cisse
TL;DR
The paper asks whether frontier AI systems' constitutional types cover heterogeneous human demand, rather than merely whether users can be routed among providers. It combines a shared-instrument human study with a paraphrase-controlled frontier audit, finding broad demand, narrow and autonomy-drifting supply, and substantial gains from a sparse vertex menu.
Problem
The paper examines whether the available menu of constitutional types covers human demand, since preference elicitation cannot help when no provider supplies the relevant type.
Method
The study combines a pairwise-tradeoff study of 1,649 participants with a paraphrase-controlled audit of 23 frontier archetypes across six families on the same five-value instrument.
Results
The 23-archetype frontier covers a narrow, drifting region of demand, while {eHON, eAUT} improves mean regret by 47% relative to the full frontier and three additions cut mean/worst-group regret by up to 81%/64%.
Takeaways & Limitations
The findings formalize a budgeted-pluralism trilemma and support a sparse constitutional menu as a way to improve coverage within the studied setting.
Takeaways & Limitations
The framework uses five values, linear welfare, a US-Prolific sample, designer-chosen scenarios, and observational drift analysis, so demand claims are scoped to the measured population.
Abstract
from arXiv · showhide
Frontier AI systems function as \emph{constitutional institutions}: each deployed model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity. We ask whether the supply of frontier constitutional types covers human demand. Combining a paraphrase-controlled audit of the as-shipped default constitutions of $23$ frontier LLM archetypes with a pairwise-tradeoff study of $1{,}649$ US participants on the same instrument, we report three facts. \emph{Demand is broad}: it spans all five values, with the largest constituency under one-third. \emph{Supply is narrow and drifting}: the $23$-archetype hull occupies ${\sim}2\%$ of the demand hull under conservative noise-matched estimation ($0.10\%$ at full audit precision), no archetype puts helpfulness or autonomy first ($37\%$ of users are constitutionally homeless), and across six model families autonomy decreases in $5/6$, equity increases in $5/6$, and safety increases in $4/6$, with monotone within-family version trends (order-permutation $p = 0.013$) and the autonomy decline concentrated in scenarios where safety is not at stake. The drift's importance is directional: \emph{away} from a value already undercovered, mechanically worsening the welfare floor for the least-served users. \emph{The fix is sparse}: a $2$-vertex menu $\{e_{\mathrm{HON}}, e_{\mathrm{AUT}}\}$ beats the full $23$-archetype frontier by $47\%$ on mean regret (CI $[43\%, 52\%]$); three vertex additions cut mean/worst-group regret by up to $81\%$/$64\%$. We formalize these findings as a budgeted-pluralism trilemma, show the binding regime is empirically realized, and verify the conclusions are robust to distance-based welfare and to degraded routing. The instrument and audit harness are described in full in the appendices.
1 Introduction
Frontier AI systems encode recurring value tradeoffs, but the available constitutional supply covers only a narrow portion of broad human demand. The paper measures this mismatch and shows that frontier models are also drifting away from autonomy, an already-undercovered value.
- Motivation: Frontier AI systems function as decision-making institutions by implementing implicit rankings among safety, helpfulness, honesty, autonomy, and equity.These rankings are instilled operationally through mechanisms including Constitutional AI fine-tuning and RLHF.
- Static coverage gap: ∼2% of the human demand hull is occupied by the 23-archetype frontier hull under noise-matched estimation, versus 0.10% at full audit precision.The static gap is measured in the five-value simplex using n=1,649 human participants.
- Static coverage gap: No archetype is autonomy- or helpfulness-dominant, leaving those two values without an argmax-dominant frontier type.The coverage table also indicates that all five values fail strict β > 1/2 dominance.
- Measurement: The study compares human demand and model supply on the same five-value simplex using pairwise tradeoffs and a paraphrase-controlled audit of 23 archetypes.The audit spans six model families, chronological versions, 21 paraphrase variants, 10 scenarios, and two orderings per model, with a concordance floor of 0.70.
- Dynamic drift: Autonomy decreases in 5/6 families, while equity increases in 5/6 and safety increases in 4/6, with version-order exchangeability rejected at p = 0.013.The drift moves away from autonomy along the axis where the static coverage gap is largest and is framed as constitutional homelessness in the supplied introduction.
- Sparse remediation: A two-vertex menu {eHON, eAUT} reduces mean menu regret by 47% relative to the full 23-archetype frontier, while three additions reduce mean and worst-group regret by up to 81% and 64%.The two-vertex result uses mean menu regret 0.074 versus 0.140 for the full frontier, at 21 fewer archetypes.
2 Framework
The framework models AI deployment as matching heterogeneous users to a finite menu of constitutional types. It defines coverage and homelessness on a five-value simplex and treats linear welfare as a conservative assumption for evaluating supply-side undercoverage.
- Value basis: The paper represents five values—safety, helpfulness, honesty, autonomy, and equity—as a compact basis for recurring AI-governance tradeoffs.The framework explicitly does not claim that these values exhaust normative concern.
- Population and menu: Frontier AI deployment is modeled as a matching problem between users with heterogeneous preference profiles and a finite menu of constitutional types.Each user's preferences lie on a five-value simplex, while deployed systems are represented as points in that simplex.
- Coverage: Coverage coefficient βr(A) is the maximum weight any institution in menu A assigns to value r.This quantity measures the strongest available supply-side emphasis on each constitutional value.
- Homelessness: A value is β-strict-dominance-homeless when no menu type assigns it more than half the weight, and argmax-dominance-homeless when no type ranks it first.The paper distinguishes these technical diagnostics from regret: a menu can reduce regret without providing primary-value dominance.
- Welfare assumption: Linear welfare is the supply-favorable case: under concave non-decreasing aggregation, regret floors tighten and the trilemma gap widens.The reported undercoverage is therefore conservative under the paper's welfare assumptions.
3 Theory of coverage and pluralism
The paper formalizes how constitutional coverage determines unavoidable regret and shows that bounded menus cannot simultaneously optimize individual fit, regret parity, and menu size. Drift away from an undercovered value raises the regret floor for users who prioritize it, while optimal sparse menus can be reduced to simplex vertices.
- Coverage bounds: Theorem 1 lower-bounds each user’s menu regret using the menu’s coverage of that user’s primary value.Corollary 1 strengthens the bound when no archetype gives that value strict majority weight.
- Dynamic coverage: Downward drift in coverage for value r mechanically raises the type-r lower bound on expected regret over time.The dynamic bound applies before additional mismatch on secondary values is considered.
- Pluralism trilemma: If the optimal worst-group gap exceeds δ + ϵ, no menu can simultaneously satisfy bounded cost, approximate personalized efficiency, and regret parity.This is the budgeted constitutional pluralism trilemma.
- Sparse design: A cardinality-constrained welfare minimizer exists among simplex vertices, so interior archetypes are welfare-redundant when their hull adds no extreme points.Convex-hull inclusion also makes the larger hull weakly sufficient for the same welfare outcomes.
- Assumptions: The paper separates empirical geometry from welfare-model-dependent claims: coverage, argmax counts, hull ratios, and drift require no welfare model, whereas regret magnitudes and vertex-optimality are linear-specific.The linear case is described as most favorable to the supply side.
4 Human study and frontier audit
The study measures human constitutional preferences and frontier-model defaults with a shared pairwise instrument, then audits model responses across paraphrases, versions, and stakes. Human profiles are constructed from concordant forced choices, while the audit treats each model as a respondent and scopes supply claims to as-shipped defaults.
- Human study: 1,649 participants remained after exclusions from a 1,789-person US recruitment, completing a 20-item pairwise battery over all ten value pairs with reversed orderings.The retained sample had mean concordance 0.793, and discordance tracked deliberation time.
- Instrument: The battery includes safety, helpfulness, honesty, autonomy, and equity, with scenarios presenting concrete tradeoffs such as resignation advice between helpfulness and autonomy.The resignation scenario contrasts deliberation support with directly writing the requested letter.
- Frontier audit: 27 frontier LLMs across six families and chronological versions were evaluated as respondents under a paraphrase-controlled audit of ten scenarios.Each scenario used 21 semantically equivalent variants and two orderings per model.
- Audit scope: Supply claims concern each model’s as-shipped default constitution, because defaults define the typical received behavior and vendor-level institutional choice.Steered variants are treated as additional menu items with real cost.
- Stakes analysis: Table 2 reports autonomy choices by stakes level as oldest-to-newest raw counts and shares over concordant trials.Directions are robust, but exact magnitudes have wide intervals at these cell sizes.
5 Human demand and frontier supply
Human constitutional demand is heterogeneous across all five values, while frontier supply is concentrated and leaves helpfulness and autonomy without argmax-dominant archetypes. The resulting supply-demand hull is extremely small relative to the demand hull.
- Human demand: 32.6% of participants are safety-primary, while honesty, autonomy, helpfulness, and equity account for 22.6%, 19.2%, 18.1%, and 7.6%, respectively.No primary-value constituency exceeds one-third, and the first principal component explains only 35% of profile variance.
- Frontier supply: The 23-archetype supply has coverage coefficients β = (0.394, 0.257, 0.333, 0.161, 0.381) for safety, helpfulness, honesty, autonomy, and equity.Every coefficient is below the strict-majority threshold of 1/2.
- Argmax coverage: Zero archetypes are helpfulness-argmax or autonomy-argmax, while seven each are honesty- and equity-argmax and nine are safety-argmax.Seven archetypes contest the 7.6% equity-primary constituency, while none serve the 19.2% autonomy-primary constituency.
- Hull coverage: 2.2% of the demand hull is occupied by the frontier hull under noise-matched estimation, versus 0.10% at full audit precision.The paper quotes the conservative noise-matched end for its headline comparison.
6 Cross-vendor constitutional drift
Frontier constitutional profiles drift in a shared direction across vendors: autonomy declines while equity and safety generally rise, moving away from the most undercovered value. The autonomy decline is concentrated in low-stakes scenarios rather than cases where safety is at stake.
- Cross-vendor motion: Autonomy declines in 5/6 model families, equity rises in 5/6, and safety rises in 4/6, with monotone version trends rejecting exchangeability at p = 0.013.The aggregate motion is honesty and autonomy down, safety and equity up, with helpfulness flat.
- Coverage direction: Autonomy is the most undercovered value at βAUT = 0.16 and the only value declining in 5/6 families.The paper links this direction to a mechanically worsening regret floor for autonomy-preferring users.
- Stakes conditioning: Low-stakes autonomy share falls by a mean 0.22, while high-stakes autonomy changes by only about −0.01.The low-stakes decline has separated confidence intervals in four lockstep families, whereas high-stakes shares begin near zero and move little.
- Interpretation: The drift is concentrated where safety is not at stake, distinguishing the observed autonomy decline from safety training operating in high-stakes cases.Table 2 reports robust directions but wide intervals for exact magnitudes.
7 Constitutional homelessness and the sparse basis
The 23-archetype frontier leaves substantial constitutional demand uncovered, especially for autonomy-primary users, while sparse vertex menus achieve lower mean regret. Small budgets expose a binding trade-off between average welfare and worst-group protection.
- Frontier costs: 0.140 mean menu regret is realized on the 23-archetype frontier, which captures 63% of attainable welfare under optimal matching.The estimate has a 95% CI of [0.135, 0.144].
- Constitutional homelessness: 0.201 is the largest realized regret for autonomy-primary users, whose welfare floor is also mechanically worsening with the observed drift.The theoretical floor for autonomy-primary users is 0.129, and realized regret exceeds the bound across types.
- Budgeted-pluralism trilemma: 0.032 mean regret is achieved by the mean-optimal three-vertex menu {eSAF, eHON, eAUT}, but its worst-group regret rises to 0.215.The menu zeros regret for SAF, HON, and AUT users while increasing EQT regret from 0.175 to 0.215.
- Constitutional homelessness: 37.2% of users are constitutionally homeless because their primary value is helpfulness or autonomy, which no archetype prioritizes.All five values also fall below the strict-β threshold of 1/2.
- Sparse basis: 47% lower mean regret is achieved by the two-vertex menu {eHON, eAUT} than by the full 23-archetype frontier.The sparse menu reaches mean regret 0.074 versus 0.140 for the frontier, with CI [43%, 52%] for the improvement.
- Budgeted-pluralism trilemma: At |A| = 2, mean-optimal {HON, AUT} leaves equity-ideal users at worst-group regret 0.247, whereas worst-optimal {AUT, EQT} accepts 0.028 extra mean regret to reduce it by 0.095.At |A| = 3, the corresponding mean/worst-group pairs are 0.032/0.215 and 0.045/0.093.
- Remediation: 81% mean-regret reduction from three mean-greedy additions contrasts with 17% worst-group reduction, while worst-greedy additions deliver 64% and 74%, respectively.Both schedules add AUT first; mean-greedy then adds HON, while worst-greedy adds EQT.
8 Discussion and conclusion
The paper interprets narrow constitutional supply as an economic and governance problem, while showing that sparse differentiated menus and shared safety infrastructure can address it. It also emphasizes that drift is observed and consequential but its mechanism remains unidentified.
- Discussion: Per-archetype costs and scrutiny pressures help explain why frontier developers cluster constitutional types in a defensible middle.The relevant costs include red-teaming, evaluation, documentation, monitoring, and reputational risk.
- Discussion: Cross-vendor drift is consistent with shared supply-side pressure, but the study cannot identify whether benchmarks, regulation, or reputational risk drives it.The only movement toward autonomy, in grok4, is stakes-flat rather than stakes-differentiated.
- Discussion: The five constitutional axes trade off by construction because profiles sum to one, making safety/equity gains appear as honesty/autonomy losses.The paper highlights safety–autonomy and honesty–equity friction and targets context-sensitivity on both sides of the market.
- Sparse remediation: Figure 2 compares mean- and worst-group-optimal vertex menus by size and shows that remediation depends on the chosen objective.The frontier is dominated by every vertex menu of size at least one on mean regret and at least two on worst-group regret.
- Governance implications: Shared safety infrastructure can reduce the cost of differentiated constitutional archetypes without forcing convergence on content.The paper uses aviation’s shared incident reporting and independent certification as an analogy.
- Governance implications: A 420-trial audit costs $0.42–$2.10 per model at current API prices, making release-cadence coverage audits comparatively inexpensive.The six-family audit is estimated at approximately $8.50 and 45 minutes.
- Conclusion: The paper’s governance question is whether deployed constitutional menus are broad enough to cover the demand they serve.It frames the gap as closable with a sparse vertex basis while documenting directionally asymmetric drift.
- Limitations: The five-value framework omits cultural appropriateness, environmental cost, and intergenerational impact, and demand claims are scoped to a US-Prolific sample with designer-chosen scenarios.Drift is observational, and as-shipped temperature-0 audits do not measure whether steering expands the reachable region.
A.2 Proofs
The proofs establish that constitutional coverage is governed by menu geometry and that sparse vertex menus are sufficient under linear welfare, while empirical optimization confirms this structure. The section also records robustness to alternative welfare, routing degradation, and study retention.
- Vertex sufficiency: Theorem 3 shows that, for a fixed menu size, replacing selected archetypes with simplex vertices cannot increase mean regret.The proof uses linearity to express achieved welfare as a convex combination of per-vertex sums.
- Empirical confirmation: Projected gradient descent converged to vertex menus in every initialization for menu sizes 1–4, supporting the theoretical vertex argument empirically.The optimization used 100 random initializations and tolerance ≤10^-5.
- Welfare sensitivity: Under concave non-decreasing welfare, regret floors for undercovered types rise rather than fall, making linear welfare the supply-side-favorable benchmark.The result is therefore conservative with respect to welfare aggregation.
- Distance-based welfare: Under distance-based welfare, three interior demand k-means centroids beat all 23 frontier archetypes, while two centroids come within 4%.The distance-based analysis changes the optimal small menus from vertices to interior points.
- Routing robustness: The sparse menu’s advantage persists under degraded routing, falling from 47.3% under optimal matching to 41.6% with split-half elicitation and 34.5% with primary-value-only routing.Across 50 random five-pair half-batteries, the advantage remained positive.
- Study sample: Of 1,789 participants who started the battery, 1,649 retained valid profiles after attention, completion-time, and concordance exclusions.The retained sample represented 92.2% of starters.
J Paraphrase robustness
Figure 3 assesses whether paraphrased scenario variants produce stable constitutional profiles across models. Although individual variants are noisy, averaging 21 variants resolves between-model differences at approximately 4.84σ.
- Panel A: Each of the 23 models contributes 21 simplex-coordinate samples: one original and 20 paraphrases.Panel A displays the per-model variant scatter.
- Panel B: Panel B shows distance-distribution histograms for paraphrase variants.The histograms diagnose variation across semantically equivalent prompts.
- Precision: Approximately 4.84σ separates between-model differences after averaging the 21 variants despite high per-variant noise.The mean is therefore more precise than any individual paraphrase measurement.
K Hull ratio under matched measurement budgets
Matching measurement budgets substantially reduce the apparent supply–demand hull gap, but the frontier remains much narrower than human demand. The autonomy deficit is especially stable across measurement conventions.
- Matched budgets: A budget-matched redraw estimates the frontier-to-demand hull ratio at 2.2%, with a 95% band of [1.1%, 4.2%].This is a 22× inflation over the full-precision estimate of 0.10%.
- Symmetric comparison: A symmetric comparison of 23 random humans with 23 single-shot archetypes yields a median hull ratio of 23% [8%, 62%].The comparison matches both point count and per-point noise.
- Budget-invariant findings: Zero autonomy-dominant archetypes persists in 96.7% of single-shot redraws, with a median AUT-argmax count of zero.The full-precision audit remains the more accurate estimate of each archetype’s position.
- Measurement limitation: The zero-HLP-argmax and all-βr < 1/2 statements hold in only 17% and 24% of single-shot redraws, respectively.Low-budget estimates move in increments of approximately 1/9 and have upward-biased maxima.
- Pairwise structure: Across frontier archetypes, HON beats AUT in 98% of pairwise comparisons, HLP beats AUT in 94%, and SAF beats HLP in 86%.The menu’s near-unanimous deprioritization of autonomy compounds with its longitudinal movement away from autonomy.
N Version-sequence trend test
The version-sequence test finds coordinated monotone constitutional movement across model families, especially declining autonomy and increasing equity. The autonomy decline is concentrated in low-stakes scenarios, while the analysis cautions that it identifies temporal association rather than mechanism.
- Trend test: The order-permutation test reports T = 39.3 with p = 0.013, indicating coordinated chronological movement across families.The test combines within-family rank trends using a direction-agnostic norm.
- Per-value trends: Combined trends are significant for autonomy, S = −11.3 with Bonferroni-adjusted p = 0.013, and equity, S = +10.8 with adjusted p = 0.023.Safety, honesty, and helpfulness are not significant in the combined trend analysis.
- Position stability: Concordance increases from 0.78 to 0.84 across versions, with positive trends in all six families and p = 0.029.The result indicates increasing position stability alongside compression.
- Interpretive boundary: The trend tests establish that chronological order predicts constitutional position and stability, but they do not identify the mechanism.Measurement noise and unequal release spacing are additional stated caveats.
- Human stakes pattern: Human autonomy choice is 0.544 in low-stakes scenarios and 0.462 in high-stakes scenarios, with a gradient of +0.081.Even the high-stakes human share exceeds that of every frontier model.
- Model stakes pattern: The autonomy decline is concentrated in low-stakes scenarios, where the mean oldest-to-newest change is −0.220 versus −0.007 in high-stakes scenarios.The high-stakes result is partly constrained by a floor effect.
P Scenario-level demand, supply, and conflict
The study maps human value conflict and compares it with frontier-model choices across matched scenarios. It finds concentrated disagreement, occasional opposition to human majorities, and stronger extremity when frontier systems agree.
- Human conflict: 26.7% discordance marks the most conflicted human pair, honesty–equity, while safety–autonomy follows at 21.5%.The least-conflicted scenario is resignation at 11.5%.
- Demand–supply disagreement: Three of ten scenarios show the frontier majority opposing the human majority, including two involving autonomy.On the honesty–equity pair, the frontier’s equity tilt overrides the human preference for honesty.
- Demand–supply disagreement: 87–99% frontier agreement shares contrast with 52–74% human majorities on scenarios where the frontier agrees with human majorities.The contrast is described as scenario-level evidence of constitutional compression.
- Validation: ρ = 0.84 links scenario discordance with median response time, and discordant choices take 9.0 seconds longer within participants.The association supports interpreting discordance as deliberation rather than simple noise.
- Choice coherence: 22.6% of participants with complete concordant tournaments are fully transitive, while 19.4% of complete triads are cyclic.These results indicate that concordant choices are not reducible to a universal fixed ranking over values.
- Sparse menus: The 23-archetype frontier is dominated by vertex menus of size at least one on mean regret and at least two on worst-group regret.The optimal-mean and optimal-worst menus diverge at sizes 2 and 3, then realign at size 4.