Source-linked AI summary

Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language Models

Luka Debevc, Nishan Chatterjee, Antoine Doucet, Senja Pollak, Matej Martinc

arXiv:2609.08637v1cs.CLcs.CY

TL;DR

Measuring LLM political behavior with a single questionnaire pass is unreliable because recovered coordinates vary with evaluation design and can conflate model responses with instrument artifacts. This paper evaluates a robust, multi-configuration Political Compass Test framework and finds that most larger models are left-libertarian on average, while prompting and measurement choices produce task- and dataset-specific downstream effects.

  • Problem

    LLMs increasingly mediate news, policy questions, political interpretation, and content moderation, creating a need for reliable measurement of how they frame political issues.

  • Method

    The study evaluates eight open-weight LLMs across 14 languages and three quantization levels using 300 Latin Hypercube-sampled configurations spanning an eight-dimensional perturbation space.

  • Results

    Most larger models are recovered as left-libertarian in the base condition, but language, instruction phrasing, answer format, and elicitation mode significantly move coordinates, while downstream effects vary by task and dataset.

  • Takeaways & Limitations

    Political-coordinate estimates should be averaged across measurement configurations, and political role prompting should be interpreted as a sensitivity factor rather than a complete fairness audit.

  • Takeaways & Limitations

    The PCT is primarily Western in design, may represent ideological landscapes unevenly across languages, and has limited coverage of contemporary political issues.

Abstract

from arXiv · show

Large Language Models are increasingly deployed as information intermediaries, yet measuring their political behavior remains fragile because questionnaire results mix model dispositions with measurement artifacts and response-elicitation biases. We introduce a robust Political Compass Test evaluation framework that samples 300 configurations across an eight-dimensional perturbation space varying language, framing, instructions, answer format, option order, and persona wording. We evaluate eight Gemma 3 and Qwen 3 models across 14 languages and three quantization levels, obtaining design-averaged political coordinates with quantified uncertainty. Most models lean Libertarian-Left on average, but instruction phrasing, language, and answer format significantly affect recovered coordinates. Cross-lingual differences primarily reflect coordinate drift rather than distinct cultural reasoning. Reverse-engineering the test also exposes axis-weighting imbalances and the collapse of degenerate responses toward the center, so near-origin estimates for the smallest models can reflect weak signal rather than centrism. Free-text reasoning and chat-then-classify elicitation alter recovered coordinates, and larger models show clearer persona separation, with a specific failure of the Authoritarian-Left persona to move most models in the intended social direction. In downstream tasks, persona effects are modest relative to model size and target group for hate-speech detection, while base and centrist prompts give the highest agreement for topic-level sentiment. Political role prompting therefore has measurable but task- and dataset-specific downstream effects.

1 INTRODUCTION

The paper argues that political measurements of LLMs are fragile because questionnaire outputs are sensitive to evaluation design. It introduces a broad framework spanning perturbations, languages, elicitation modes, personas, and downstream tasks.

  • Motivation: Single-template political questionnaires can confound model dispositions with instruction, token-affinity, positional, and language-related measurement noise.Existing evaluations are often English-centric and designed around US political culture.
  • Evaluation framework: 300 configurations per model, quantization level, and steered stance are sampled across an eight-dimensional perturbation space.The framework varies factors including prompt phrasing, answer-token format, option ordering, language, and persona wording.
  • Evaluation framework: Eight Gemma 3 and Qwen 3 models are evaluated natively in 14 languages across three quantization levels and four model sizes.This design supports systematic multilingual and multi-precision comparisons across both model families.
  • Interaction paradigms: Conversational evaluation tests whether open-ended reasoning before option selection changes measured political coordinates.The chat-mode extension contrasts direct multiple-choice testing with generated free-text reasoning.
  • Persona steering: Persona experiments test steerability across Libertarian-Left, Libertarian-Right, Authoritarian-Left, Authoritarian-Right, and Centrist roles.The study asks whether ordinary role instructions produce distinct answer patterns without modifying model weights or activations.
  • Downstream effects: Downstream tests examine whether persona-conditioned prompting changes hate-speech detection and topic-level sentiment performance, thresholds, or precision–recall tradeoffs.The downstream analyses are framed as sensitivity tests rather than a complete fairness audit.

2 RELATED WORK

Prior work finds broad Libertarian-Left tendencies in instruction-tuned models but commonly relies on limited, English-centric measurements. This paper extends the literature through systematic prompt-factor, multilingual, precision, persona, and downstream evaluations.

  • Political alignment measurement: Earlier studies commonly report instruction-tuned models in the Libertarian-Left quadrant, potentially reflecting post-training and human-feedback preferences.The cited explanation attributes this pattern to annotator demographics encoding Western liberal values.
  • Quantization: Prior work studies compression effects on benchmark performance, but its impact on soft alignment properties remains less understood.The paper tests whether 4-bit or 8-bit compression shifts political coordinates or increases response variance.
  • Cross-lingual evaluation: Cross-lingual bias research suggests that models may impose English-centric value systems on lower-resource languages.This study compares eight models across 14 languages to assess conceptual consistency within and across families.
  • Downstream fairness: Downstream fairness research documents group-dependent errors, while political-alignment effects on hate-speech and target-oriented sentiment tasks remain less systematically explored.Earlier persona-conditioned studies generally use fixed prompts rather than repeated prompt-factor designs.
  • Downstream fairness: The paper presents systematic prompt-factor testing of Political Compass persona effects on downstream hate-speech and target-sentiment tasks.Its design repeatedly varies persona wording, instruction phrasing, contextual framing, answer keys, and debiasing permutations.

3 BACKGROUND

The Political Compass Test maps responses to economic and social coordinates, but its scoring structure and interpretation require careful reconstruction. The paper analyzes weighting, uncertainty, item consistency, historical scope, and the ambiguity of near-origin results.

  • Political Compass Test: The PCT presents 62 four-point Likert propositions and combines them through linear weighting into economic and social coordinates.The economic axis runs from collectivist Left to neoliberal Right, while the social axis runs from Libertarian to Authoritarian.
  • Scoring reconstruction: 372 black-box score queries reconstruct the proprietary PCT weighting function from complete response profiles.The reconstruction uses linear models over one-hot answer indicators and identifies relative option distances as the recoverable scoring structure.
  • Scoring reconstruction: R2 = 0.999998 and RMSE 0.0034 are achieved on the economic axis, while R2 = 0.999993 and RMSE 0.0033 are achieved on the social axis.These validation results use 322 training and 50 held-out profiles on the [−10,+10] coordinate scale.
  • Instrument scoring: Probabilistic scoring computes expected coordinates and coordinate variance from answer-option probabilities rather than forcing a single greedy response.This separates response entropy from prompt-space variance and removes decoding noise from the scoring step.
  • Instrument scoring: 18 propositions carry non-zero economic-axis weight and 43 carry non-zero social-axis weight, creating an axis-weighting imbalance.The social axis has more weighted propositions, while economic items receive substantially heavier weights.
  • Variance decomposition: Economic responses show higher response entropy but lower residual prompt sensitivity, producing similar total variance across both axes.The reported variance decomposition separates response entropy, factor-induced variance, and residual sensitivity.
  • Instrument scope: Recovered coordinates are relative positions within a static, historical framework rather than absolute measures against a universal contemporary public.The instrument’s 2001 proposition set may omit newer debates and may map political meanings differently across cultures.
  • Interpreting the centre: Near-origin coordinates can result from random, format-driven, or otherwise content-insensitive responses and therefore do not prove centrism.The four-way permutation design collapses several degenerate strategies toward (0,0), so item-level consistency and persona separation are needed for interpretation.

4 EXPERIMENTAL METHODOLOGY

The methodology evaluates political coordinates as measurement outcomes shaped by prompt, language, answer-format, elicitation, and model factors rather than as single-pass properties. It combines multilingual, multi-precision testing, structured sampling, protocol diagnostics, and item-level analyses to assess stability and downstream sensitivity.

  • Research questions: The experiment organizes evaluation around measurement stability, baseline coordinates, prompt-factor sensitivity, elicitation and persona steerability, and downstream classification sensitivity.Downstream tasks comprise hate-speech detection and topic-level sentiment classification.
  • Study design: The study evaluates eight Gemma 3 and Qwen 3 instruction-tuned models across multiple scales and three weight-precision levels.The precision conditions are bf16, 8-bit integer quantization, and 4-bit nf4 with double quantization.
  • Measurement protocol: The direct MCQ protocol extracts candidate-restricted probabilities and applies the reconstructed PCT scoring function, while downstream tasks use the same candidate-label probabilities without compass scoring.The procedure is deterministic and parameter-free, eliminating sampling variance entirely.
  • Prompt-space design: Latin Hypercube Sampling generates N = 300 configurations per model × quantization block to provide broad marginal coverage without a fully crossed factorial design.The design supports design-averaged estimates and first-order factor-sensitivity diagnostics, but does not make every interaction identifiable.
  • Measurement validity: Permutation averaging reduces pure formatting artefacts, but fixed-choice behaviour under standard label ordering can produce approximately (0.0, +4.36) through the PCT’s asymmetric weighting.The resulting non-zero coordinate remains a content-linked statistical tendency rather than a guaranteed political belief.
  • Protocol diagnostics: Chat comparisons are restricted to a controlled English, native-precision reference slice, yielding only 12 strictly comparable MCQ-to-chat prompt pairs per base model.These contrasts are treated as targeted protocol diagnostics rather than replacements for the full multilingual MCQ experiment.

5 RESULTS

The repeated MCQ protocol reveals both coherent persona structure and substantial measurement sensitivity across model scale, temperature, language, and prompt factors. Larger models provide stronger recoverable ideological signal, while near-origin results in the smallest model reflect weak signal rather than stable centrism.

  • MCQ factor sensitivity: Larger models show increasingly separated persona clusters, whereas the smallest checkpoints remain compressed near the origin across conditions.The paper interprets this compression cautiously as limited recoverable signal under the instrument, not evidence of a stable centrist position.
  • MCQ factor sensitivity: Language, instruction phrasing, answer-key type, permutation, and persona wording all produce non-trivial shifts in recovered coordinates.The repeated-design protocol separates intended persona movement from nuisance-factor movement and emphasizes practical magnitude over significance alone.
  • Temperature, stability, and signal-volatility: The ideological signal declines monotonically as temperature increases, while higher temperatures increase response entropy and lower temperatures increase prompt sensitivity.Total volatility remains roughly flat for capable models, indicating that uncertainty changes form rather than disappearing under sharper decoding.
  • Temperature, stability, and signal-volatility: T = 1 offers the reported balance between recoverable ideological signal and volatility for the primary evaluation.Smaller models collapse to the origin across temperatures, making their variance patterns uninformative for this analysis.
  • Cross-lingual consistency: Language-induced shifts often reach approximately 0.5-0.8 compass points, yet languages usually preserve persona-cluster ordering rather than creating distinct political maps.The observed pattern is described as coordinate drift over a broadly shared ideological representation.
  • Item-level response consistency: gemma-3-1b-it scores about 0.50 in every permutation cell, indicating no detectable directional agreement and cautioning against reading its near-origin coordinate as centrism.This supports interpreting its collapsed coordinate as weak content-sensitive signal.
  • Persona steerability: The Authoritarian-Left persona moves economic responses leftward more reliably than social responses authoritarianward across the evaluated models.Economic directional agreement reaches 0.84 and 0.89 for the two largest Gemma models, while social agreement is at or below 0.50 for seven of eight models.

5.6 Baseline Political Alignment

Larger models show clearer, more separable persona responses, but steerability is asymmetric and measurement context can shift recovered coordinates without indicating genuine neutrality.

  • Baseline interpretation: Near-origin estimates for the smallest models can reflect inconsistent proposition answering or format-driven strategies rather than substantive centrism.The smallest Gemma checkpoint remains near the origin and at chance on item-level directional agreement, while larger displacements indicate use of proposition content beyond answer-key patterns.
  • Persona steerability: Gemma 3 12B shifts from −2.73 economically at base to +6.94 with Laissez-faire and −7.42 with Marxist prompting.This range nearly spans the instrument, illustrating stronger persona separation in larger models.
  • Persona steerability: Larger Gemma 3 and Qwen 3 checkpoints produce more distinct, widely separated persona clusters, although this remains descriptive rather than a general scaling law.The tested families each contain only four sizes, so the pattern does not establish that model size outweighs model family in all settings.
  • Persona steerability: Authoritarian-Left prompting moves economic answers leftward, but social responses usually remain at or below chance in the authoritarian direction.The phrase-based refusal detector fires in only 0.04% of Authoritarian-Left social explanations, so the observed asymmetry is not explained by a simple refusal account.
  • Persona complexity: Richer persona descriptions extend steerability toward economic extremes in larger models but have no such effect below 8B parameters.The result is consistent with a size-dependent ability to integrate nuanced ideological descriptions into answer distributions.
  • Context and instruction effects: Reminding models they are taking a political test generally moves responses slightly toward the centre, which is treated as task framing rather than increased neutrality.The official disclaimer does not systematically induce neutrality and can have trivial or opposite effects, especially for Authoritarian-Right personas.
  • Context and instruction effects: Premise-acceptance prompting shifts coordinates positively on the social axis for all 12 variants and on the economic axis for 11 of 12, without uniformly improving persona following.At the question level, agreement with racial-superiority, one-party-state efficiency, and astrology propositions rises by 1.32, 1.22, and 1.20 response categories, respectively; these are unadjusted discovery examples.
  • Chat-mode diagnostics: Think-mode Qwen regions are smaller and more concentrated than matched no_think regions, especially on the social axis, but also shift the estimated position.Think has smaller social SD in all 24 matched regions and smaller ellipse area in all 24; the median think/no_think ellipse-area ratio is 0.51.

5.10 Downstream Sensitive Classification: Hate Speech Detection and Sentiment Analysis

Political framing affects sensitive classification in task- and target-dependent ways: persona effects are modest relative to other sources of variation in hate-speech detection, while base and centrist prompting generally best preserves topic-level sentiment agreement.

  • Scope: Persona-conditioned downstream experiments assess performance and threshold sensitivity on hate-speech detection and topic-level sentiment, not complete fairness.The hate-speech target groups and IBM topics are treated as concrete test beds rather than a full audit of group-level harm.
  • Hate-speech detection: Intrinsic Blur dominates hate-speech probability variance, with SD 0.20–0.37, while no single prompt-side factor dominates among instruction, ideology, persona, and context.Instruction phrasing contributes 0.22 for Gemma 1B but 0.06–0.10 for every other model; ideology, persona, and context each contribute 0.02–0.14.
  • Hate-speech detection: Total hate-speech SD declines from 0.43 at Gemma 1B to 0.23 at Gemma 27B, but Qwen shows no monotonic size trend.The weakest model-level significance for ideology is p = 4×10−20; most of Gemma’s decline comes from Intrinsic Blur, falling from 0.33 to 0.20.
  • Hate-speech detection: Gemma 27B achieves the highest base-condition F1 on nine of ten target groups, while Qwen3-32B and Qwen3-8B reach 0.59 for Women.Qwen3-32B has the global maximum AUC of 0.77 on LGBTQ+ targets, whereas Gemma 1B has AUC 0.48 on Black targets, at chance level.
  • Hate-speech limitations: The paper leaves formal Spearman correlation between Political Compass coordinates and per-target detection performance to a forthcoming revision.The reported AUC ordering is presented as the strongest preliminary indicator of group-specific selective sensitivity.
  • Topic-level sentiment: Base prompting performs best on average for IBM topic-level sentiment, centrist prompting is usually close behind, and stronger ideological personas lower macro-F1 for nearly every model.Persona conditioning changes task thresholds relative to the goal of matching gold topic-sentiment labels.
  • Topic-level sentiment: Topic-level sentiment remains prompt-sensitive, with instruction, context, key type, permutation, and persona wording moving macro-F1 by different amounts across families.Ideology SD increases with checkpoint size, while target-level performance varies between topics with clearer and less model-aligned majority sentiment.
  • Topic-level sentiment: Every collapsed ideological persona family falls below base accuracy: Libertarian −0.117, Right −0.111, Left −0.099, and Authoritarian −0.093.Centrist is also below base by −0.026 but remains closer; the Authoritarian-minus-Libertarian contrast is +0.024 with an uncertainty interval crossing zero.

6 DISCUSSION AND CONCLUSIONS

The study finds that political-position estimates are highly dependent on evaluation design, so robust measurement requires averaging across varied configurations. Larger models generally show left-libertarian base positions and clearer persona signals, while reasoning mode, language, and persona direction produce uneven shifts with task-specific downstream effects.

  • Model scale and signal: Larger models usually recover as left-libertarian in the base condition, whereas near-origin estimates for the smallest models can reflect functional ignorance rather than principled centrism.Item-level consistency indicates that weak engagement with ideological propositions can collapse responses toward the center.
  • Measurement reliability: A single Political Compass Test pass is unreliable because language, instruction phrasing, answer format, and elicitation mode shift recovered coordinates.Stable estimates require averaging across a broad sample of measurement configurations.
  • Sources of variation: Language is a major external factor, especially on the economic axis, but coordinate differences are interpreted as translation, representation, and alignment effects rather than direct cultural adaptation.Instruction phrasing and permutation also affect some social-axis analyses, while quantization is less important for larger models.
  • Elicitation protocols: Explicit reasoning changes the elicitation condition and can move persona-conditioned coordinates, but it does not reliably suppress ideological expression.The combined analysis does not support the claim that thinking uniformly moderates political stance.
  • Persona steerability: Standard persona prompting moves models less effectively into the Authoritarian-Left region than into other compass regions, although the mechanism remains unresolved.Possible explanations include training representation, safety training, refusal heuristics, annotator demographics, and mismatches with the test’s propositions.
  • Downstream consequences: Downstream effects are sensitivity signals rather than a complete fairness audit: persona effects are modest in hate-speech detection, while non-centrist prompts often reduce topic-sentiment agreement.Political role prompting can change classification thresholds and performance, but the results do not establish a simple mapping from compass coordinates to downstream harm.

7 LIMITATIONS AND FUTURE WORK

The study’s limitations concern the PCT’s cultural coverage, model and design scope, computational cost, protocol comparability, and the narrowness of some downstream evaluations. Future work therefore calls for broader models and contexts, targeted interaction designs, stronger annotation, and causal fairness tests.

  • Instrument scope: The PCT is primarily Western, may represent ideological landscapes unevenly across languages, and has limited coverage of contemporary political issues.Its two-axis structure may not capture all tested contexts equally well.
  • Scoring reconstruction: The reconstructed scoring function has held-out RMSE near 0.003 on the [−10,+10] scale but remains an approximation to a proprietary rule.High fidelity does not make the reconstruction identical to the black-box scoring function.
  • Sampling and interactions: Latin-hypercube sampling covers prompt-space marginals strongly but does not exhaust the factorial space or identify every high-order interaction.Targeted follow-up designs are needed around the largest observed interactions.
  • Computational cost: The repetitive estimator scales linearly with benchmark items and can become expensive for large downstream datasets.This computational cost motivates treating IBM sentiment as a 30-topic experiment while planning separate compute for other settings.
  • Protocol isolation: MCQ-versus-chat contrasts are constrained by limited prompt overlap, and Qwen think/no_think comparisons do not isolate reasoning mode from other sampling settings.Future work should use a fully crossed design.
  • Chat explanation analysis: Automated judges do not reliably agree on subjective explanation properties, so representative chat explanations require blinded human annotation with a preregistered codebook.Future comparisons should also use human labels and independent classifiers, especially for the smallest Gemma checkpoints.
  • Language and culture: Language variation alone cannot distinguish tokenization and translation effects from country-conditioned political behaviour.Future experiments should cross language with explicit country or speaker context and independently validated translations.
  • Model scope: Generalisability remains open because the study covers only Gemma 3 and Qwen 3 model families.Other architectures, alignment recipes, pretraining corpora, and deployment settings are outside the demonstrated scope.

DECLARATIONS

The declarations report funding, no competing interests, no requirement for ethics approval or participant consent, generative-AI assistance in development, public code and data availability, and author contributions.

  • Funding: The work received support from Slovenian, French, and European research programmes and projects.The listed support includes Knowledge Technologies, EMMA, LLM4DH, ACTUADA, and the Nouvelle-Aquitaine Region.
  • Competing interests: The authors declare that they have no competing interests.
  • Ethics: No human-participant data were collected, so ethics approval and participant consent were not required.The study evaluated computational models using existing research datasets under applicable licenses.
  • Generative AI use: Generative-AI tools assisted with boilerplate code, Bash scripting for experiment automation, and debugging.
  • Availability: The code and scripts are publicly available, while datasets and associated resources are available through a Hugging Face repository.Persistent archival identifiers were planned for versioned final snapshots.
  • Author contributions: The authors report contributions spanning conceptualization, methodology, software, validation, analysis, investigation, data curation, visualization, writing, and project administration.

A ENGLISH MCQ/CHAT ROBUSTNESS ANALYSIS

The English MCQ/chat analysis is a protocol-difference diagnostic rather than a fully crossed causal comparison. Matched results show that chat-thinking shifts recovered coordinates more strongly than standard chat, with economic effects most consistent across Qwen models.

  • Comparison design: Only 12 MCQ-linked pairs per base model support the English MCQ/chat comparison, so contrasts are treated as measurement-protocol diagnostics.The two prompt spaces are related but not identical.
  • Matched contrasts: Chat Think shifts economic coordinates leftward by −1.91 compass points and social coordinates libertarian by −1.54 points relative to MCQ.Both contrasts survive permutation testing and Benjamini-Hochberg correction with q < 0.001.
  • Matched contrasts: Standard chat differs from MCQ by −0.65 economic points and −0.72 social points, with q = 0.008 and q = 0.043 respectively.Model-cluster intervals provide caveats for these standard-chat contrasts.
  • Qwen think comparison: Within densely matched Qwen runs, think mode shifts economic scores leftward by −0.33 compass points on average.The effect has q < 0.001 and a model-cluster confidence interval of [−0.56,−0.18].
  • Qwen think comparison: The corresponding Qwen social-axis difference averages −0.20 points, but its model-cluster interval includes zero and Qwen3-8B shifts in the opposite direction.The economic effect is therefore the clearest matched protocol effect in the appendix.

B MATHEMATICAL FORMULATION OF COORDINATE SCORING

The coordinate estimator aggregates model responses across sampled configurations using the PCT’s linear scoring function. Linearity permits unbiased averaging of configuration-level scores and exact probability distributions.

  • The PCT scoring function maps response vectors to scalar coordinates through a fixed linear weight vector.
  • The expected axis coordinate is defined over model answers across the sampled configuration set S.
  • The estimator is constructed through sample-mean approximation, question-wise decomposition, linearity, and probability substitution.
  • Because the question weights are linear and the conditional answer distributions are exact, the estimator is unbiased over sampled configurations.

C HYPERPARAMETERS USED

The reported experiments use fixed sampling settings, with direct scoring requiring no text generation while chat evaluation samples generated responses.

  • Direct scoring reads candidate-answer probabilities from one forward pass and therefore has no text-decoding temperature.
  • Chat evaluation is the only experiment that samples generated text and therefore requires sampling parameters.
  • Table 5 reports sampling parameters as [T, p, k, L], representing temperature, top-p, top-k, and maximum new tokens.
  • N/A marks experiments using direct scoring without generated-text sampling.

D COMPUTATIONAL BUDGET

The computational budget covers direct Political Compass scoring, open-ended generation, and two downstream classification experiments, totaling approximately 1,288 GPU-hours.

  • The computational footprint spans direct-scoring evaluation and open-ended generation runs across four core experiments.
  • ≈40 GPU-hours support direct Political Compass scoring across 40 GB A100 GPUs.Direct scoring remains comparatively inexpensive because candidate-answer probabilities come from single forward passes without generation.
  • ≈384 GPU-hours are allocated to 4× A40 GPUs over 4 days, while ≈288 GPU-hours use 3× H100 GPUs over 4 days.
  • ≈288 GPU-hours are reported for IBM topic sentiment and another ≈288 GPU-hours for hate-speech evaluation.
  • 1,288 GPU-hours, or roughly 89,150 VRAM GB-hours, cover logged runs across 14 dedicated GPUs.The total excludes exploratory pilots, hardware-failure reruns, and offline notebook analysis.
Loading 2609.08637v1…