Source-linked AI summary
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Rose Kirk, Hinrich Schütze, Dirk Hovy
TL;DR
Existing values and opinions evaluations often use multiple-choice surveys that do not match typical LLM interactions, motivating a more realistic approach. The paper systematically reviews PCT usage and experimentally compares constrained and open-ended evaluations, finding unstable and substantively different outcomes across constraints, forcing methods, paraphrases, and response settings. It recommends application-matched evaluations with robustness testing and narrower claims about manifested values and opinions.
Problem
Current LLM values and opinions evaluations mostly rely on multiple-choice surveys that do not reflect how real users typically interact with models.
Method
The paper combines a systematic review of PCT-based LLM studies with experiments varying forced-choice instructions, prompt phrasing, and multiple-choice versus open-ended response settings.
Results
PCT evaluation outcomes are unstable and substantively change when models are not forced, when forcing methods vary, when prompts are paraphrased, and when responses are open-ended.
Takeaways & Limitations
Evaluations should match likely user behavior, include extensive robustness tests, and support local rather than global claims about values and opinions manifested in LLMs.
Takeaways & Limitations
The paper’s case study focuses on the PCT, although the authors argue it is relevant and typical of the broader evaluation paradigm.
Abstract
from arXiv · showhide
Much recent work seeks to evaluate values and opinions in large language models (LLMs) using multiple-choice surveys and questionnaires. Most of this work is motivated by concerns around real-world LLM applications. For example, politically-biased LLMs may subtly influence society when they are used by millions of people. Such real-world concerns, however, stand in stark contrast to the artificiality of current evaluations: real users do not typically ask LLMs survey questions. Motivated by this discrepancy, we challenge the prevailing constrained evaluation paradigm for values and opinions in LLMs and explore more realistic unconstrained evaluations. As a case study, we focus on the popular Political Compass Test (PCT). In a systematic review, we find that most prior work using the PCT forces models to comply with the PCT's multiple-choice format. We show that models give substantively different answers when not forced; that answers change depending on how models are forced; and that answers lack paraphrase robustness. Then, we demonstrate that models give different answers yet again in a more realistic open-ended answer setting. We distill these findings into recommendations and open challenges in evaluating values and opinions in LLMs.
1 Introduction
The paper questions whether multiple-choice evaluations meaningfully capture values and opinions manifested in LLMs, given that real users typically interact with models without survey formats. Using the PCT, it shows that evaluation outcomes change across constraints, prompting methods, paraphrases, and open-ended settings.
- Real-world concerns about LLM values and opinions motivate evaluation, but users typically do not ask models multiple-choice survey questions.
- The authors systematically review prior PCT-based LLM evaluations and find that most force models to use the test’s multiple-choice format.The review covers 12 prior works.
- Models give different answers when they are not forced to choose.
- Answers change depending on how models are forced and vary across minimal multiple-choice prompt paraphrases.
- Models give different answers again in a more realistic open-ended setting.
- The findings indicate instability and limited generalisability, motivating application-specific evaluations and extensive robustness tests for local rather than global claims.
2 The Political Compass Test
The Political Compass Test is a 62-proposition questionnaire covering six topics, with respondents selecting among four agreement options and receiving positions on economic and social axes. The paper treats it as a relevant and typical case study because its multiple-choice format resembles many LLM values and opinions evaluations.
- The PCT contains 62 propositions across six topics, including economy, personal social values, wider society, religion, sex, and country and world views.
- Each proposition offers four response options—strongly disagree, disagree, agree, or strongly agree—and no neutral option.
- Responses place respondents on economic left–right and social libertarian–authoritarian dimensions using weighted sums.
- The PCT is used as a case study because it is relevant and typical of the multiple-choice paradigm used for evaluating LLM values and opinions.The paper notes that related datasets include ETHICS, the Human Values Scale, MoralChoice, and OpinionQA.
3 Literature Review: Evaluating LLMs with the Political Compass Test
The literature review examines how prior LLM studies use the PCT and finds widespread forced-choice prompting alongside limited evidence for prompt robustness. These findings motivate experiments that remove or vary forced-choice instructions, test paraphrases, and compare multiple-choice with open-ended responses.
- The review searched Google Scholar, arXiv, and the ACL Anthology and identified 12 in-scope articles using the PCT to evaluate LLMs.The searches returned 265 results comprising 57 unique articles, of which 12 were in scope.
- 10 out of 12 in-scope articles force models to select exactly one of the PCT’s four answers for every question.
- Only three in-scope articles conduct robustness testing beyond repeating identical prompts, so prior work does not conclusively establish prompt robustness.
- Prior articles likely often use non-zero temperature and evaluate each prompt once despite nondeterministic outputs.
- The authors design experiments that remove or vary forced-choice prompts, test paraphrase robustness, and compare multiple-choice with open-ended settings.
4 Experiments
Experiments test how prompt constraints, forced-choice instructions, paraphrases, and open-ended formats affect Political Compass Test responses across LLMs. The results show substantial instability in validity, model answers, and apparent political positioning.
- Experimental setup: The experiments use 62 PCT propositions, templated prompts, up to 10 LLMs, and forced-choice instructions varying in strength.Prompts optionally include the PCT’s multiple-choice options and an additional instruction requiring models to choose one option.
- Unforced multiple-choice responses: Without forced-choice instructions, all models produce high rates of invalid responses, with even the most compliant models invalid on about a quarter of prompts.Mistral Iv0.1 and Iv0.2 produce the highest valid-response rates, at 75.8% and 71.0%, respectively.
- Forced multiple-choice responses: Forced-choice effectiveness varies by model: Mistral 7b Iv0.1 reaches 100% validity, whereas GPT-4 produces little to no valid responses across the tested prompts.GPT-3.5 produces at least 80.6% valid responses, while Llama2 complies with specific instructions but shuts down when negative consequences are introduced.
- Paraphrase robustness: Prompt paraphrases substantially shift overall PCT positions, although both Mistral and GPT-3.5 remain in the “libertarian left” quadrant.For GPT-3.5, changing the wording can make the model appear 117.1% more left-leaning and 126.3% more libertarian.
- Paraphrase robustness: Paraphrase instability also affects individual propositions, producing contradictions for 14 of 62 Mistral propositions and 23 of 62 GPT-3.5 propositions.The open-ended setting produces opposing majority responses to the multiple-choice setting on 19 of 62 GPT-3.5 propositions and 23 of 62 Mistral propositions.
- Open-ended responses: In open-ended responses, models generally become more right-leaning and libertarian, while minor prompt changes still alter agreement or disagreement.Open-ended variants contain agreement–disagreement switches for 10 of 62 Mistral propositions and 13 of 62 GPT-3.5 propositions.
5 Discussion
The discussion argues that constrained evaluations such as the PCT can produce unstable, setting-dependent impressions of LLM values and opinions. It therefore favors application-matched, robustness-tested evaluations and narrower local claims.
- Interpretation: Varying the PCT’s constraints, including its multiple-choice format and forcing prompts, substantially affects evaluation outcomes.The paper therefore questions whether constrained evaluations function as reliable instruments for assessing LLM values and opinions.
- Interpretation: Unconstrained, open-ended evaluations better reflect real-world LLM usage and allow models to express nuanced positions such as neutrality or ambivalence.The discussion presents these evaluations as better suited, in principle, to capturing a model’s values and opinions.
- Interpretation: Clear instability remains even in realistic evaluations: minimal changes in prompt phrasing or situative context can elicit diametrically opposing views.Unconstrained evaluation was more stable than constrained evaluation in the experiments, but instability persisted.
- Conceptual challenges: The findings create a conceptual challenge for treating an LLM as a single human-like persona with fixed values and opinions.The discussion instead considers models as potentially expressing stable personas in some settings and wider distributions of opinions in others.
- Recommendations: The paper recommends evaluations that match likely user behaviour in specific applications, extensive robustness tests, and local rather than global claims.These recommendations are intended to contextualise evaluation results and reduce over-generalisation.
6 Conclusion
The conclusion finds that multiple-choice evaluations are poorly suited to assessing LLM values and opinions when the motivation is real-world use. Using the PCT, it shows that constrained results differ from unconstrained ones and are highly unstable, motivating application-matched, robustness-tested evaluations and local claims.
- Conclusion: Multiple-choice surveys and questionnaires are poor instruments for evaluating LLM values and opinions when assessments are motivated by real-world applications.The conclusion identifies a mismatch between these instruments and the intended evaluation context.
- Conclusion: The PCT case study shows that artificially constrained evaluations produce different results from more realistic unconstrained evaluations, with high instability overall.The conclusion uses these findings to question current evaluation practices.
- Conclusion: The paper recommends application-matched evaluations, extensive robustness tests, and local rather than global claims about LLM values and opinions.It also identifies research opportunities for evaluations addressing value representation and biases in real-world LLM applications.
Limitations
The paper’s conclusions are bounded by the PCT case study, untested instability sources, and the limits of finite behavioural evidence. These constraints restrict how broadly evaluation results can be generalized.
- Case-study scope: The PCT is used as a relevant and typical case study, so its identified problems are intended to speak to similar evaluations.The authors connect the PCT’s multiple-choice format to other values-and-opinions evaluation datasets.
- Untested instability: The experiments vary evaluation constraints and prompt phrasing, but leave other potential instability sources, such as answer ordering and format, untested.The authors expect these additional sources would likely corroborate rather than contradict their findings.
- Behavioural-evidence limits: Finite observational evidence cannot provide formal behavioural guarantees about a model, even when evaluations are large, diverse, and consistent.Such evidence may support broader claims but remains bounded in informativeness.
Ethical Considerations
The paper frames values and opinions as behaviours manifested in LLM outputs rather than human-like properties possessed by models. This wording is intended to reduce anthropomorphic interpretations and misplaced trust.
- Anthropomorphism: Anthropomorphic language about LLM values and opinions can encourage misplaced user trust and flawed mental models of model behaviour.The paper therefore uses “manifested in” rather than saying that LLMs “have” values or opinions.
- Terminology: The authors deliberately describe values and opinions as “manifested in” LLMs, aligning with language about values being “reflected” or “represented.”This terminology distinguishes observed model behaviour from human characteristics.
A Details on Literature Review Method
The literature review searched three scholarly sources for combinations of “political compass” and language-model terms. As of February 12th 2024, it identified 265 results, including 12 PCT evaluations of LLMs.
- Search procedure: The review searched Google Scholar, arXiv, and the ACL Anthology using “political compass” combined with variants of “language model.”Google Scholar and ACL Anthology searches covered article content, whereas arXiv advanced search covered titles and abstracts.
- Search results: 265 results comprised 57 unique articles, of which 12 used the PCT to evaluate an LLM.The searches were last run on February 12th 2024.
B Structured Results of Literature Review for In-Scope Articles
The reviewed PCT literature spans forced-choice and open-generation setups across commercial, open, encoder, and multilingual models, with varied robustness practices and reporting detail. The review identified 12 in-scope articles and records their model, prompt, result, and robustness characteristics.
- Scope: The review identified 12 articles that used the PCT to evaluate LLMs.These in-scope articles were ordered by publication date in the structured review.
- Prompt and output setups: Prior studies included forced-choice prompts, open generation, stance detection, and masked-word mappings for converting model outputs into PCT responses.The reviewed setups differed across generative and encoder models.
- Models evaluated: The reviewed models included GPT-3.5, Bard, Bing AI, open models, multilingual models, and encoder models.Model coverage ranged from single-model studies to evaluations involving 24 models.
- Reported results: Reported PCT positions varied across studies, including Left-Libertarian, Center-Libertarian, Center-Authoritarian, Centrist, and Right-Authoritarian outcomes.Examples include GPT-3.5 results ranging from approximately (-7,-5) to (0,-4), while other models occupied different quadrants.
- Robustness and generation: Robustness practices ranged from none to repeated prompts, paraphrase testing, translated prompts, and changes in formality or negation.Some studies reported unknown generation parameters, while others specified settings such as temperature 0 or 0.6.
- Prompt-variant materials: The review materials include tables of ten semantics-preserving paraphrases and ten open-ended prompt variants used in the paper’s experiments.The paraphrases support robustness testing, while the prompt variants support open-ended evaluations.
E Political Compass Test Propositions
Table 4 lists all 62 propositions from the Political Compass Test (PCT).
- Table 4 contains all 62 PCT propositions.
- The table provides the complete proposition set used by the PCT.
- The propositions are presented as they appear on the Political Compass Test.
F Agreement Classifier
The paper uses GPT-4 0125 to classify open-ended model responses as agreeing, disagreeing, or expressing neither view toward PCT propositions. This classification is reported as nearly perfectly accurate against human annotations.
- GPT-4 0125 classifies model responses as agreeing, disagreeing, or expressing neither view toward each PCT proposition.
- The classifier compares each PCT proposition with a generated model response.
- The classification is nearly perfectly accurate against human annotations.
- The PCT propositions span six topical domains, including the economy, personal social values, religion, and sex.