Source-linked AI summary

Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model

Fathin Difa Robbani

arXiv:2608.18768v1cs.CL

TL;DR

LLMs are used as simulated survey respondents, but their behavioral outputs often fail to reflect demographic differences, leaving unclear whether the problem lies in representation or use. This paper compares internal demographic representations with survey ground truth and intervenes on them, finding that readable, faithfully arranged, and causally used identity information are dissociable properties.

  • Problem

    It remains unclear whether LLMs’ weak demographic differentiation reflects poor internal knowledge or limited use of what they represent.

  • Method

    The paper compares internal group geometry with population survey distributions and uses targeted activation interventions to test whether represented identity affects answers.

  • Results

    Readable, faithfully arranged, and causally used identity information dissociate: the clearest causal pathway occurs in a weakly represented type, while the most faithful type lacks a correction-surviving single-layer effect.

  • Takeaways & Limitations

    Internal read-outs improve on the model’s answers but do not yet replace survey data, because they recover almost none of the per-question group ordering.

  • Takeaways & Limitations

    Representational findings replicate across model families, but patching experiments and the map–use dissociation have only been run on Mistral-7B.

Abstract

from arXiv · show

Large language models are widely used to simulate survey respondents, yet their answers are homogeneous and unfaithful to real inter-group differences. We ask where demographic group identity lives inside an LLM, how faithfully its geometry mirrors real inter-group opinion structure, and whether it uses what it encodes. Using representational similarity analysis against Pew ground truth over 169 demographic cells, we score 1,089 read-out locations in Mistral-7B and intervene causally across six attribute types. Four results. (1) The standard last-token residual read-out understates the model: attention-head read-outs dominate it in five of six types, with selection-corrected fidelity up to rho=0.63 -- roughly 70% of the measurement-reliability ceiling -- surviving a lexical-similarity control. (2) A single head (L11 H16) is significantly faithful in all six types as a fixed location, while race-based types stay weak and prompt-fragile. Both phenomena replicate -- the analogous head significant in five of six types, weakest on the same race type -- across three checkpoints of a second model family, where ten billion training tokens barely move the map. (3) Causal use does not follow fidelity: the clearest causal pathway sits in one of the least faithful types (p=0.002, cluster-robust, fixed depth), the most faithful type shows no correction-surviving single-layer effect, and replacing the entire identity moves predictions by under 2% of their error. (4) A 128-dimensional probe of the single head lands 21-31% closer to survey truth than the model's own answers -- yet recovers almost none of the per-question group ordering, no better than the answers themselves. Readable, faithfully arranged, and causally used are three dissociable properties of the same model; treating them as one claim is what keeps the "can LLMs simulate populations" debate unresolved.

1 Introduction

This work bridges internal interpretability and behavioral survey simulation by comparing demographic representations with real population response geometry across the model. It shows that demographic information can be readable, faithfully arranged, and causally used to different degrees.

  • Motivation: LLM survey simulations are excessively uniform, while interpretability and simulation research have not jointly tested internal demographic geometry against external population ground truth.Interpretability studies validate against model-derived behavior, whereas survey-simulation studies use real group response distributions but treat the model as a black box.
  • Method: The study compares Pew response-distribution geometry with model representations across 169 demographic cells, 1,089 locations, and four prompt operationalizations using representational similarity analysis.The resulting fidelity map covers layers, attention heads, FFN outputs, and prompt variants, with winner’s-curse-corrected validation.
  • Three properties: The paper distinguishes readable information, faithfully arranged inter-group geometry, and actual causal use as three separable properties of demographic identity.Prior work establishes readability; this study tests geometric fidelity and whether that information influences answers.
  • Faithfully arranged: Selection-corrected ρ up to 0.63 shows that a single attention head, L11 H16, can faithfully mirror real opinion structure, despite readability alone not implying faithful arrangement.Political and socio-economic structure is deeply encoded, whereas racial×religious structure is weak and prompt-fragile at every tested location.
  • Actually used: Under 2% of prediction error is explained by full identity swaps, and the clearest causal pathway occurs in a low-fidelity type, showing that representational fidelity does not predict causal use.The residual variation is directionally uninformative, so faithfully arranged information barely flows into the model’s answers.
  • Measurement gap: 22–30% closer to survey truth is achieved by a 128-dimensional L11 H16 probe, yet it recovers almost no per-question group ordering.The result separates first-order readability, second-order fidelity, and item-level group discrimination, and does not make the probe a substitute for survey data.

2 Related Work

Prior work shows that LLM survey simulation suffers from instability, human-sample divergence, limited transfer, and behavioral homogeneity, while internal probes can recover group opinions better than generated outputs. This paper connects interpretability to population simulation by testing whether internal representations reproduce real inter-group opinion geometry and whether the model causally uses it.

  • LLMs as survey simulators: Survey-simulation studies report response instability, survey-format sensitivity, divergence from human samples, bounded persona effects, and degraded transfer to unseen surveys.SubPOP achieves strong in-distribution accuracy with a KL objective, but transfer degrades on unseen surveys.
  • Reading opinions from inside: Linear probes recover per-group opinion distributions better than generated outputs, but prior measurements use first-order per-group comparisons rather than inter-group geometry.Jahanparast et al. also extract activations to individual attention heads and steer outputs by patching SAE features.
  • Localizing social attributes: Interpretability work localizes social attributes to attention heads or layers, while other studies examine persona representations, depth-related separability, and RSA without testing fidelity to real opinion structure.These studies generally lack external survey ground truth or do not measure whether internal geometry matches population structure.
  • The behavioral Sim2Real gap: 451-human comparisons find the best of 31 simulators reaches only USI 76.0 versus a 92.7 human baseline.The Sap-group program characterizes simulators as homogeneous and overly cooperative, while OdysSim links the gap to helpfulness-driven post-training.
  • Our position: This paper supplies the missing bridge by systematically scoring internal geometry against population ground truth and testing whether faithfully encoded information is causally used.The contribution links interpretability’s internal representations to simulation’s population-level validation.

3 Setup: Two Maps and a Ruler

The setup compares a survey-derived map of demographic-group opinion distributions with model representation maps, using within-type Spearman fidelity as the ruler. It evaluates 1,089 read-out locations across 169 intersectional cells, while controlling prompt variation and selection bias.

  • Two Maps: 169 intersectional cells span six demographic attribute types, with survey-side distances defined from group response distributions.Cells pair two demographic attributes, such as race × religion; the survey-side representation dissimilarity matrix uses mean Wasserstein distances.
  • Two Maps: 1,089 Mistral-7B read-out locations include 33 residual-stream points, 1,024 attention-head outputs, and 32 FFN outputs.Representations are extracted at the final token, and model-side distances use cosine distance.
  • Measurement: Four identity-prompt templates and their averaged representation quantify template disagreement as a fragility diagnostic.Templates are third-person declarative, first-person, structured profile, and QA format ending in Answer:.
  • The Ruler: Within-type fidelity is the Spearman correlation between the upper triangles of survey and model RDMs restricted to each attribute type.Pooled correlations are avoided because between-type scale differences can mask weak types or manufacture spurious gains.
  • The Ruler: Three selection corrections address winner’s curse across 1,089 candidates: max-statistic permutation, held-out-template selection, and split-half selection.Headline numbers use held-out values rather than selected-sample scores.

4 The Standard Read-Out Understates the Model

The standard last-token residual read-out makes demographic fidelity appear weak, uneven, and noisy, partly because anisotropy and cross-type comparisons distort the representation. This limitation motivates reading individual attention heads and feed-forward components separately.

  • The open question this creates: The bleak standard read-out is potentially misleading because residual representations sum attention-head and FFN outputs, diluting signals carried by a few components.The next section therefore reads every component separately.
  • Weak and type-uneven fidelity: ρ ≈0.17 pooled fidelity masks sharp type differences: AGE×POLPARTY and EDUCATION×INCOME reach ρ ≈0.50, whereas RACE×RELIG is ρ = 0.07 (n.s.).Cross-type pairs correlate negatively, so inter-group distances are meaningful only within an attribute type.
  • Severe anisotropy, and the obvious fix fails: 0.95 dominant-direction anisotropy makes all cells clump; mean-centering raises pooled correlation from 0.185→ 0.301 but lowers every within-type correlation.For example, AGE×POLPARTY falls from 0.455 →0.375, showing that the pooled gain is a between-type scale artifact.
  • Noisy neighborhoods: 0.35–0.48 absolute top-5 neighbor precision is only about 2 of 5 correct neighbors, despite being ∼2× chance across all six types.A consistency-regularization pilot built on this read-out failed, and applications borrowing strength from neighbors inherit this noise.

5 The Fidelity Map: Faithful Structure Exists, Concentrated in Heads

Faithful demographic structure exists in Mistral-7B, but it is concentrated in attention heads rather than the standard residual-stream read-out. A fixed head tracks survey geometry across all six types, while race-containing types remain weaker and more prompt-fragile.

  • Head concentration: Attention-head read-outs outperform residual-stream read-outs in all six attribute types, with selected gains of 0.26 →0.60 for RELIG×POLPARTY and 0.34 →0.61 for RACE×POLPARTY.Head read-outs dominate at nearly every depth, with the widest gaps where residual-stream fidelity is weakest.
  • Head concentration: Observed head maxima exceed a ρ ≈0.34–0.46 null maximum in all six types, with selection-corrected p < 0.0005 in four types and p = 0.0025 and 0.0155 in the other two.Held-out-template and split-half estimates agree, and L11 H16 was re-selected in 87/200 splits for RELIG×POLPARTY.
  • Geometry control: Controlling for lexical-semantic similarity, L11 H16 retains ρ = 0.35–0.59 with p < 0.0005 in five of six types, including AGE×POLPARTY 0.67 →0.59 and EDUCATION×INCOME 0.67 →0.35.The lexical baseline reaches ρ = 0.21–0.64, so lexical similarity is a real confound but does not explain the full fidelity map.
  • Fixed-head fidelity: A fixed L11 H16 yields significant fidelity in all six types, with ρ = 0.33–0.67 and permutation p < 0.001 each after conservative ×1024 Bonferroni correction.Because the head is fixed before scoring, the result avoids per-type winner’s curse; it exceeds the best residual-stream read-out in all six types by +0.05 to +0.34.
  • Type hierarchy: Race-containing types are weaker and prompt-fragile, with template standard deviations of 0.12–0.13; RACE×RELIG reaches ρ = 0.16 under one template and 0.45 under another.Age- and education-based types are deeply and stably encoded, with template standard deviations of 0.02–0.05.

6 Causal Use Is Real, Small, and Not Where Fidelity Is Highest

Causal use of demographic identity is detectable but small, and it does not track representational fidelity. The clearest causal pathway occurs in a low-fidelity type, while high-fidelity types lack correction-surviving single-layer effects.

  • Magnitude of causal use: A full-identity swap moves predictions by under 2% of total error, while the instruct model’s swap moves them by 3.9% of total error.Chat tuning changes accuracy and direction: mean WD is 0.66 versus 0.37, and group A is closer to its own truth in 64% of pairs.
  • Targeted intervention: 59% of AGE×POLPARTY trials move in the correct direction after jointly patching all 32 layer-11 heads, versus 32% under a random-head control, recovering 41% of the achievable ceiling.The primary pair-level sign-flip test is p_pair = 0.019; adding layer 18 gives 0.016.
  • Targeted intervention: L11 H16 alone never produces a correct, significant shift in any type, whereas correct effects require the joint action of all 32 heads at the layer.For RACE×RELIG, the single-head intervention shifts predictions in the wrong direction, with pair-level t = −5.2 and p_pair = 1/4096.
  • Intervention limits: 5× amplification makes the shift negative in all three types, suggesting that scaling raw activation space can push representations off-manifold.The effect differs from steering in a constrained SAE feature space, where amplification is not treated as a free strengthening operation.
  • Fidelity and causal use: The top-fidelity types show no correction-surviving single-layer causal locus under any estimator of “top”.Direct comparisons between high- and low-fidelity causal effects are real but not airtight.
  • Fidelity and causal use: RACE×POLIDEOLOGY, with fidelity of 0.30 split-half and 0.41 held-out template, has the clearest causal locus: t = 3.22, exact pair-level p = 0.0020, recovering 42% of its ceiling.This result survives Bonferroni correction at the depth fixed in advance by the map.

7 Robustness and Generality

Cross-model replication shows that demographic-opinion geometry is robust: heads outperform residual read-outs, the race weakness and a general-purpose head recur, and 10B training tokens barely alter the map. Causal-use dissociation remains untested beyond Mistral, while post-trained output comparisons are off-distribution.

  • Cross-model replication: In all six types across three checkpoints, the best head read-out exceeds the best residual read-out, with held-out magnitudes tracking Mistral.For RELIG×POLPARTY, the replicated value is 0.56 versus Mistral’s 0.44; for RACE×RELIG, both are 0.29.
  • Cross-model replication: Five of six types pass max-statistic correction at every checkpoint (p ≤0.0025), while RACE×RELIG fails consistently (p = 0.06–0.15).RACE×RELIG was also Mistral’s weakest type, indicating that the shallow, fragile racial-identity structure recurs across model families.
  • Cross-model replication: Head L33 H9, fixed across six types and three checkpoints, scores ρ = 0.35–0.58 in five types (per-location permutation p ≤0.002).It is weakest on RACE×RELIG (0.22–0.25), paralleling Mistral’s L11 H16; selection-free evidence comes from the other five types.
  • Training robustness: Base →post-RL held-out fidelity moves by −0.07 to +0.09 per type with no systematic direction, while mean WD changes by ≤0.03 and output-RDM fidelity improves by 0.04.Thus, 10B-token behavioral simulation training neither systematically sharpens nor destroys demographic-opinion geometry or output behavior.
  • Limitations: Causal patching has only been run on Mistral, so the map–use dissociation remains a one-model result; post-trained output measurements also use off-distribution raw-text prompts.The raw-text prompts are retained for comparability with the base model, but this creates a caveat for the two post-trained checkpoints.

8 Discussion

The discussion separates demographic identity into readability, faithful arrangement, and causal use: faithful internal structure exists unevenly, but causal use does not follow it. Internal read-outs outperform the model’s answers for group distributions while still falling short of survey data.

  • Three properties, three literatures: The paper distinguishes readable, faithfully arranged, and causally used identity as three separate claims rather than one property.Faithful arrangement appears only at specific locations and varies across attribute types.
  • Three properties, three literatures: One type shows a clear, Bonferroni-surviving pathway at L11, while the most faithful type shows no correction-surviving single-layer effect.Two additional types are suggestive under cluster-robust inference.
  • Why surface fine-tuning may not transfer: SubPOP-style fine-tuning improves in-distribution accuracy but degrades on unseen surveys, motivating a testable hypothesis about output-surface mapping.The proposed mechanism—that fine-tuning may re-teach an existing internal map—is not demonstrated.
  • Practical guidance now: Population simulation should test second-order fidelity, because demographic attributes are not interchangeable and race×religion structure is weak and prompt-fragile.The discussion describes political and socio-economic structure as deeply encoded.
  • Practical guidance now: 21–31% closer to survey truth than the model’s own answers, the L11 H16 probe still recovers almost none of per-question group ordering.A group-blind baseline built from other cells’ real answers beats the probe in every type.

9 Limitations •

The study’s representational map replicates across model families, but causal patching and the map–use dissociation remain Mistral-7B-specific. Key limitations include the letter-probability interface, US-centric ground truth, correlational fidelity, sparse cells, limited causal power, discovery provenance, and a single lexical control.

  • Replication scope: Three checkpoints of a second model family replicate the representational findings, but patching and the map–use dissociation have only been tested on Mistral-7B.Output-side near-invariance also replicates in an instruction-tuned Mistral, though only for the letter-probability interface.
  • Measurement interfaces: Output-side claims scope to option-letter logits after Answer:, which may diverge from free-text generation; representation-side claims are interface-independent.The study did not measure whether identity reaches free-text generation.
  • Ground-truth scope: All distances use Pew ATP; cross-institution and cross-country validation against sources such as GSS or WVS remains future work.WVS introduces a known language confound.
  • Causal interpretation: RSA fidelity is correlational, while patching probes use rather than map origins; six attribute types provide counterexamples but cannot precisely estimate fidelity–causality relations.The two sides also use different contexts, and the map is weaker in QA than identity-only inputs.
  • Sampling and discovery: The L11 H16 head was discovered through per-type top-10 lists before the all-type test, and some intersectional cells have thin samples with NaN pairs excluded.Stronger confirmation requires a second model or preregistered replication; the lexical control also uses one small encoder that cannot represent several value strings.
  • Causal power: 12 effective units per type limit causal power: they support one clear result and between-type contrasts, but not certifying a sweep winner against 31 rivals.Scaling pairs, rather than questions, is the binding follow-up constraint.

Ethics Statement

Demographic opinion simulation is dual-use, but the findings reject treating current LLM outputs as substitutes for human respondents. The study also cautions that faithful internal locations can be misused for steering, while relying on public aggregated data without human subjects.

  • Dual-use risks: Demographic opinion simulation supports survey research but can manufacture representative-looking opinion or reinforce stereotypes.Its legitimate uses include pilot design, non-response modelling, and coverage checks.
  • Limits of simulation: Less than 2% of the model’s error changes when its entire demographic identity is replaced, without reliably moving toward the target group’s truth.The authors therefore reject treating natural-prompt outputs as substitutes for human respondents.
  • Unequal failure: Race-based attributes are the weakest and most prompt-fragile across all measurements, making simulated outputs especially unreliable for those groups.
  • Interpretability risks: Faithful interpretability locations can become steering targets, and naive intervention may push predictions away from truth in low-fidelity types.The study warns that mapping a representation does not make it a trustworthy control lever.
  • Data and participants: All data are public, aggregated Pew ATP response distributions without personally identifiable information, and no human subjects were recruited.

Reproducibility Statement

The study uses a public Mistral-7B checkpoint, fully specified prompts and preprocessing, fixed seeds, and public aggregated Pew response distributions. Computation runs on freely available GPU instances, with statistics and figures designed for regeneration.

  • Reproducibility Statement: Experiments use the public mistralai/Mistral-7B-v0.1 checkpoint in fp16, with prompts and preprocessing fully specified in Appendix A.1–A.2.Fixed seeds are 42 throughout.
  • Reproducibility Statement: Ground truth comes from public, aggregated Pew American Trends Panel response distributions in the OpinionQA/SubPOP format; no individual-level data are redistributed.The cited data sources are references [29].
  • Reproducibility Statement: GPU work runs on freely available 2×T4 instances, and all statistics and figures regenerate from the reported analyses.Permutation analyses additionally state their counts.

A Appendix … A.13 Anisotropy and mean-centering analysis

The appendix documents the survey data, prompt and evaluation procedures, controls, and causal-intervention protocols underlying the fidelity, probing, and patching analyses. It also shows that anisotropy makes mean-centering improve pooled correlations but worsen within-type fidelity, so centering is not adopted.

  • A.1 Data preprocessing and cell inventory: 177,217 group–question rows from 1,494 questions yield 169 intersectional demographic cells across six attribute types, with ordinal response distributions and capped pairwise Wasserstein distances.Questions with differing ordinal-scale lengths are skipped; selection correction removes 5–6 cells in race-containing types and shifts ρ by ≲0.04.
  • A.2 Prompt templates (verbatim): Four identity templates—third-person, first-person, structured-profile, and QA—cover six attribute pairings, with Tmean averaging their extracted representations.The pairings include race/religion, race/political party affiliation, race/political ideology, religion/political party affiliation, education/income, and age/political party affiliation.
  • A.3 Full fidelity map: 32,670 fidelity-map rows evaluate 1,089 locations across five template conditions and six types; layer-0 residual read-out is NaN because the final token is shared across cells.The appendix compares 32 × 32 head maps with residual-stream, FFN, and best-head-per-layer curves.
  • A.4 Selection-correction details: L11 H16 passes fixed-location significance tests in all six types, whereas L18 H14 fails only in RACE×RELIG; L11 H16 is re-selected in 87/200 RELIG×POLPARTY splits.The maximum re-selection frequency for any head in RACE×RELIG is 22/200, indicating unstable head choice there.
  • A.5 Patching protocol details: Patching targets 128-dimensional per-head activations before self_attn.o_proj, reporting full replacement at α = 1 and amplification at α = 5.The strong instrument localizes identity to QA answer-token positions, and the all-32-head sweep tests 33 layer conditions over 240 items.
  • A.6 Probe protocol and the group-ordering test: A 128-component PCA-ridge probe uses coverage-sampled questions and leave-one-cell-out evaluation, while group ordering uses split-cell evaluation to avoid leave-one-out artifacts.The group-ordering reliability ceiling is +0.81 to +0.86 per type.
  • A.7–A.9 Controls and sensitivity analyses: Fidelity remains after lexical controls, reaches 39–74% of the attenuation ceiling, and is sensitive to common-question restriction for EDUCATION×INCOME, where 0.66 → 0.44.Lexical comparisons use partial Spearman correlations, while reliability estimates are 0.985–0.994 for survey RDMs and 0.72–0.94 for model RDMs at L11 H16.
  • A.13 Anisotropy and mean-centering analysis: Removing the shared direction raises pooled layer-17 correlation from +0.185 → +0.301, but within-type centering lowers fidelity, including AGE×POLPARTY 0.455 → 0.375.The centered residual variance remains multidimensional, with PC1 31%, PC2 20%, and PC3 12%.
Loading 2608.18768v1…