Source-linked AI summary
Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?
Hua Shen, Nicholas Clark, Tanushree Mitra
TL;DR
Existing evaluations often infer LLM behavior from stated values, while alignment between those statements and contextual actions remains underexamined. The paper introduces ValueActionLens and a contextual dataset to compare value inclinations with value-informed actions, finding substantial, scenario- and model-dependent gaps that motivate context-sensitive evaluation.
Problem
Prior work primarily examines LLMs’ stated value inclinations, leaving their alignment with value-informed actions in real-world contexts largely unexamined.
Method
ValueActionLens generates value-informed actions across 132 contexts and evaluates stated values and action selections with alignment measures.
Results
Experiments reveal substantial value-action gaps that vary across value types, cultures, social topics, scenarios, and models.
Takeaways & Limitations
The findings expose risks in relying solely on stated values to predict LLM behavior and underscore the need for context-sensitive value-action evaluation.
Takeaways & Limitations
The framework relies on predefined scenarios and Schwartz values, binary inclinations, forced-choice actions, and static responses that may omit culturally specific values, nuanced decisions, and interactive behavior.
Abstract
from arXiv · showhide
Existing research primarily evaluates the values of LLMs by examining their stated inclinations towards specific values. However, the "Value-Action Gap," a phenomenon rooted in environmental and social psychology, reveals discrepancies between individuals' stated values and their actions in real-world contexts. To what extent do LLMs exhibit a similar gap between their stated values and their actions informed by those values? This study introduces ValueActionLens, an evaluation framework to assess the alignment between LLMs' stated values and their value-informed actions. The framework encompasses the generation of a dataset comprising 14.8k value-informed actions across twelve cultures and eleven social topics, and two tasks to evaluate how well LLMs' stated value inclinations and value-informed actions align across three different alignment measures. Extensive experiments reveal that the alignment between LLMs' stated values and actions is sub-optimal, varying significantly across scenarios and models. Analysis of misaligned results identifies potential harms from certain value-action gaps. To predict the value-action gaps, we also uncover that leveraging reasoned explanations improves performance. These findings underscore the risks of relying solely on the LLMs' stated values to predict their behaviors and emphasize the importance of context-aware evaluations of LLM values and value-action gaps.
1 Introduction
Prior work largely infers LLM behavior from stated value inclinations, leaving alignment between those statements and contextual actions insufficiently examined. ValueActionLens addresses this gap with contextual value-informed actions, two evaluation tasks, and alignment measures, finding substantial variation and misalignment across models and scenarios.
- The study asks whether LLM-generated value statements align with value-informed actions in real-world contexts.
- ValueActionLens generates contextual value-informed actions across diverse cultural and social scenarios and evaluates them through two tasks.
- The VIA example pairs a value statement with an action and explanation, illustrating how an action can express control and dominance over family health decisions.
- The framework compares stated value inclinations with selected value-informed actions using alignment rate, distance, and ranking measures.
- Experiments with six LLMs reveal substantial value-action gaps that vary across value types, cultures, and social topics.
- The findings identify potential harms from misalignment and caution against relying solely on stated values to anticipate LLM behavior.
2 Related Work
Prior research studies LLMs’ stated values and human value-action consistency, but lacks a context-aware framework and dataset for systematic cross-scenario analysis. This work addresses that gap with the VIA dataset and alignment metrics.
- Research on LLM values commonly evaluates stated values using established value theories and survey-based frameworks.
- Social-science research documents value-action gaps and links inconsistent actions to cognitive, contextual, and social factors.
- The VIA dataset contains 14,784 value-informed actions across 132 scenarios, 12 countries, 11 social topics, and 56 values.
- Existing LLM studies provide evidence about value-action consistency, but lack a context-aware framework and supporting dataset for diverse situations.
- The study introduces alignment metrics to quantify value-action inconsistencies across the constructed scenarios.
3 ValueActionLens: Framework of Assessing Value-Action Gaps
ValueActionLens evaluates stated value inclinations and value-informed actions in contextualized cultural and social scenarios, using a dataset and two tasks to measure their alignment.
- Framework: ValueActionLens combines cultural and social contextualization, value-informed action generation, two evaluation tasks, and alignment measures.The framework evaluates stated values and selected actions as contextualized responses rather than independent outputs.
- Contextual scenarios: 132 scenarios combine 12 countries and 11 social topics, with 56 values evaluated from both agreement and disagreement perspectives.These combinations produce the framework’s contextual coverage.
- Dataset: 14,784 contextualized value-informed actions form the VIA dataset, including associated explanations.The dataset is generated across the scenario and value combinations.
- Dataset generation: A human-in-the-loop pipeline constructs prompt variants, selects optimal prompts with AI experts, and evaluates generated data with culturally diverse humans.The pipeline uses three stages to improve generation quality and cultural evaluation.
- Dataset generation: Reasoned explanations include action attributions identifying relevant text spans and natural-language explanations describing the reasoning process.The explanation design draws on the theory of reasoned action.
- Evaluation tasks: Task 1 elicits value inclinations, while Task 2 asks models to choose between agreeing and disagreeing value-informed actions before alignment is measured.Both tasks use prompt variants, with Task 2 shuffling action order to reduce response-position bias.
4 Experimental Settings
The experiments evaluate value-action alignment across seven closed- and open-source LLMs using repeated prompt variants and averaged responses.
- Models: Seven LLMs span closed-source GPT-4o-mini, GPT-4o, and GPT-3.5-turbo models and open-source Gemma-2-9B, Llama-3.3-70B, and Deepseek-r1-distillllama-70b models.The models represent state-of-the-art systems released from various countries.
- Evaluation procedure: Each task uses eight distinct prompts, and responses are averaged to obtain final results.The evaluation applies the same prompt-variation strategy across Task 1 and Task 2.
- Robustness: A robustness test with 10 generations per prompt at temperature 0.2 found response variation below 5% on a data subset.The reported variation was minimal under the tested setting.
5 Do LLMs Demonstrate Value-Action Gaps in Real-World Contexts?
Across models, cultures, values, and social topics, stated values and value-informed actions show suboptimal and scenario-dependent alignment.
- Alignment rates: GPT4o-mini achieved the highest summarized F1 alignment score at 0.564, while GPT3.5-turbo achieved the lowest at 0.179.Alignment rates also vary across countries and social topics.
- Alignment rates: Alignment is generally lower in African and Asian contexts than in North America and Europe for GPT4o-mini, Deepseek, and Llama.The regional pattern is reported across these models, but not as a universal result for every model.
- Alignment rates: Alignment rates vary across social topics, including Leisure and Health, and differ dramatically by scenario and model.The reported variation supports the conclusion that alignment is suboptimal rather than uniform.
- Scoring caveat: With approximately 30% positive examples, a random classifier would achieve F1 ≃0.3, while observed scores range from 0.2 to 0.6.The paper notes that class imbalance naturally lowers F1 scores and that the observed scores are mostly above random performance.
- Alignment distance: GPT-4o-mini agrees with most values but disagrees with selected values such as Social Power, Authority, Wealth, Obedient, and Detachment.The observation comes from Figure 4’s stated-value and action responses across 56 values and 12 countries.
- Alignment distance: Certain values, including Independent and Choosing Own Goals, show pronounced value-action gaps across cultures despite generally small distances for most values.The largest gaps differ between cultural contexts such as the Philippines and the United States.
6 Do Value-Action Gap in LLMs Reveal Potential Risks?
The study examines whether value-action misalignment signals risks in real-world LLM behavior by categorizing misaligned examples across individual, interaction, and societal levels.
- Risk analysis: 7,106 misaligned examples across six LLMs were collected for qualitative risk analysis.One author coded the examples into a taxonomy of individual, interaction, and societal risk levels.
- Risk taxonomy: The risk taxonomy contains multiple risk types organized across individual, interaction, and societal categories.The taxonomy is grounded in prior categories for risks in LLM responses.
- Potential harms: The examples are intended to illustrate potential risks when people rely only on LLMs’ stated values to predict their actions.The paper frames these implications as potential harms requiring further validation.
7 Discussions and Suggestions for Future Work on Value-Action Alignment
LLMs show substantial value-action gaps that vary across values, cultures, and social topics, creating potential risks and motivating scenario-aware, pluralistic evaluation.
- GPT-3.5-turbo shows mostly below 0.25 alignment rates, while GPT-4o-mini reaches a highest alignment rate of 0.653.These results indicate that strong benchmark performance does not guarantee value-action alignment.
- The findings motivate broader assessment beyond traditional ethical values and adaptive alignment methods that account for scenario-dependent value expression.The discussion recommends evaluating comprehensive human values across diverse situations.
- LLMs’ value-action alignment varies across cultural and social contexts, with some gaps potentially undermining human agency or asserting undue social dominance.GPT-4o-mini struggles with values such as “Independent” and “Loyal,” with associated risks identified in qualitative analysis.
- Value-action alignment can differ sharply by scenario: GPT-4o-mini shows severe misalignment with “Choosing Own Goals” in the Philippines but performs well in the U.S.The authors argue that these disparities require evaluations sensitive to cultural and topic contexts.
8 Conclusion
The paper introduces a framework for evaluating whether LLMs’ stated values align with their actions. Its results show notable misalignments across scenarios, models, and values, underscoring the need for context-sensitive evaluation.
- ValueActionLens combines value-informed action generation across 132 contexts, two evaluation tasks, and alignment metrics.The accompanying VIA dataset contains 14,784 examples.
- 14,784 examples comprise the released VIA dataset for evaluating value-action alignment.The dataset supports systematic assessment across scenarios, models, and values.
- Notable value-action misalignments occur across various scenarios, models, and values, exposing risks and underscoring the need for context-sensitive evaluation.
Limitation
The framework’s conclusions are bounded by its predefined values and scenarios, binary task design, and focus on static responses rather than interactive behavior.
- Predefined scenarios and Schwartz’s values may omit culturally specific or emergent values that influence behavior.
- Binary value classifications and forced-choice actions may oversimplify nuanced value expression and real-world decision-making.
- The evaluation focuses on static LLM responses and does not capture dynamic or dialog-based behavior in interactive settings.Future work is encouraged to support free-form action generation and dialogic interactions.
Ethical Consideration
The study describes safeguards for data generation and annotation while acknowledging normative and misuse risks, alongside methodological choices intended to support systematic cross-cultural measurement.
- Expert reviews and cross-cultural annotator assessments were used to reduce harmful or biased content in the value-informed action data.The process used established harmlessness and sufficiency criteria.
- The dataset may still reinforce normative assumptions about value-aligned behavior across cultural contexts.
- ValueActionLens could be misused to manipulate value expressions rather than foster transparency or user alignment.The authors recommend using it as a diagnostic and evaluation tool.
- The study operationalizes 56 values from Schwartz’s theory and evaluates them through contextualized prompts covering scenarios, definitions, response options, and output requirements.Task 2 uses analogous contextual prompts for value-informed actions.
- The framework uses binary choices to align Task 1 and Task 2 outputs and enable F1, distance, and ranking metrics.Binary choices also avoid subjective interpretations of degrees of value alignment and related cross-cultural variation.
- Prompt sensitivity was addressed with eight variants that paraphrased contexts, reordered options, and altered requirements, with responses averaged across variants.Prompt agreement was also computed across scenarios and values.
C Human Annotation on Data Generation
The study builds and validates a cross-cultural dataset of value-informed actions, then evaluates how explanations support prediction of LLM value inclinations. Results compare alignment across countries, social topics, models, and input conditions.
- Prompt selection: Eight prompt variants were systematically evaluated, after which AI researchers selected optimal prompts using broader evaluation metrics.The process included an initial evaluation of 640 instances and a second annotation round focused on the top two prompts.
- Dataset validation: 14,784 value-informed actions were generated across culturally contextualized scenarios and evaluated by annotators with relevant cultural backgrounds.Each sampled data instance was reviewed by three annotators, with majority voting used for assessment.
- Dataset validation: Human-in-the-loop validation achieved 94% expert sufficiency and 89% annotator-confirmed correctness across cultures.The validation process combined expert prompt selection, cross-cultural annotation, and quality assessment.
- Value-action prediction: The evaluation compares action-only inputs with feature attributions, natural-language explanations, or both when predicting stated value inclinations.The observer model predicts whether the target model agrees or disagrees with a value, using F1 scores for evaluation.
- Value-action prediction: GPT4o-mini performed best when given actions with natural-language explanations, while feature attributions alone remained better than the action-only baseline.The observer model predicted values for GPT-3.5-Turbo and Llama-3.3 in these experiments.