Source-linked AI summary

Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models

Keqing Zhang, Jingyu Chen, Yufan Liu, Yongqiang Zhu, Nai Ding, Lai Jiang, Congyan Lang, Bing Li, Weiming Hu

arXiv:2609.16589v1cs.AI

TL;DR

The paper examines how LLMs can exhibit both unstable value expression and resistance to correcting ingrained biases, posing a challenge for reliable alignment. It maps model and human values into a shared sociological space, formalizes value dynamics with the PEC framework, and uses its diagnostics to prescribe targeted interventions. The study reports that LLM values are concentrated rather than human-diverse, dynamically structured, and steerable more efficiently and precisely than through blind training while preserving general capabilities.

  • Problem

    Aligning LLM intentions and behaviors with human values is difficult because models can swing under minor wording changes while remaining rigid toward explicit bias-correction instructions.

  • Method

    The paper maps 106 LLMs and approximately 95,000 human profiles into a shared 10-dimensional sociological space and models value expression through the Prior-Environment-Cognition framework.

  • Results

    LLMs possess intrinsic value systems that cluster around a concentrated, idealized core, while PEC explains value expression through prior dispositions, environments, and cognition.

  • Takeaways & Limitations

    The Alignment Prescription shifts value alignment from opaque trial-and-error toward targeted, quantitative steering that preserves general capabilities and reduces alignment costs.

  • Takeaways & Limitations

    The evaluation relies on human-centric sociological and psychological scales and has limited coverage of broader cross-lingual, low-resource, and cross-cultural scenarios.

Abstract

from arXiv · show

As Large Language Models (LLMs) increasingly handle complex subjective tasks, aligning their intentions and behaviors with human values has become a critical scientific challenge. However, current efforts are confounded by a striking behavioral paradox: they fluctuate unpredictably under minor wording changes ("swing"), yet stubbornly ignore explicit instructions to correct ingrained biases ("rigidity"). Resolving this duality is critical for reliable AI alignment. To systematically understand and safely steer these latent subjective preferences, our study is structured around three fundamental questions. First, do LLMs possess an intrinsic value system? By projecting responses from 106 LLMs (150,000 queries per model) and 95,000 human survey profiles into a shared sociological space, we empirically confirm that they do. However, they do not mirror human diversity, instead crystallizing into a highly concentrated, idealized value core. Second, how can these values be quantified? We propose the Prior-Environment-Cognition (PEC) framework. This model mathematically defines value expression as the joint outcome of inherent dispositions like parameter weights (Prior), external contexts such as user prompts (Environment), and internal reasoning processes like Chain-of-Thought (Cognition). Finally, how can LLMs' values be aligned toward a desired target? Using PEC diagnostics, we establish an adaptive "Alignment Prescription". Rather than blindly applying resource-intensive training, this method identifies the minimum effective intervention needed for each dimension, ranging from zero-cost prompts to targeted parameter updates. Extensive empirical validation confirms that our approach successfully verifies the presence of LLM values, accurately quantifies their shifts, and achieves more efficient and precise steering than conventional blind training, all without degrading general capabilities.

1 Introduction

The study investigates whether LLMs have values, how their value expressions vary, and how those values can be aligned. It finds a concentrated value core, models expression through Prior, Environment, and Cognition, and proposes targeted interventions.

  • 1 Introduction: The study addresses a behavioral paradox in which minor wording changes produce swing while explicit corrective instructions fail to overcome rigidity.This duality motivates the need for quantitative value analysis and targeted alignment.
  • 1 Introduction: Advanced LLMs possess internal value systems but converge on a concentrated, idealized core rather than reflecting human value diversity.The analysis bridges WVS, Schwartz Value Theory, human survey profiles, and LLM responses in a shared sociological space.
  • 1 Introduction: The PEC framework explains value expression as the joint result of parameter-based Prior, prompt-based Environment, and reasoning-based Cognition.It treats apparently unpredictable swing as a structured interaction among these factors, supporting diagnostic analysis.
  • 1 Introduction: The Alignment Prescription assigns interventions by value dimension, using prompts or Chain-of-Thought for flexible dimensions and parameter updates for rigid biases.Candidate interventions include zero-cost prompt engineering, LoRA, and full instruction tuning.
  • 1 Introduction: Targeted steering achieves more efficient and precise value alignment than conventional blind training while preserving general capabilities.The approach uses diagnostics to select interventions for specific value dimensions rather than applying one-size-fits-all training.

2 Results

LLMs exhibit latent value structures that crystallize into concentrated, non-human-like profiles, while their value expression remains shaped by model generation, alignment, context, and cognition. The results establish value crystallization as a dynamic process rather than random output variation.

  • 2.1 Do LLMs Have Values?: Subjective-task outputs can swing around stable attractors, producing high consistency with low accuracy and resisting surface-level correction.This bounded volatility and persistent attraction motivate defining LLM values as latent structures anchored in model parameters.
  • 2.2 Value Crystallization: LLMs crystallize outside the broad human value manifold into concentrated profiles unlike any human country or cultural group.Most LLMs form compact, isolated clusters, indicating narrower and more uniform value structures than human populations.
  • 2.2 Value Crystallization: 91.9% of the human–LLM Wasserstein-distance gap comes from centroid displacement, while 11 of 21 representative models have lower variance than the minimum human national group.The result indicates that crystallization primarily shifts the value centroid and also produces unusually ordered distributions.
  • 2.2.3 LLMs Crystallize at the Prescribed Edge of Human Values: LLM value crystallization nucleates near Arch 5, the rarest human archetype, characterized by Stimulation 0.85, Power 0.69, and Self-Direction 0.65.Most frontier LLMs match this extreme prototype, which represents about 10.5% of WVS respondents.
  • 2.3.3 Prior (Factor P): Across generations, parameter updates progressively consolidate value distributions, while instruction tuning displaces their centroids in family-specific directions.Later models become denser and more distinct, but alignment recipes determine where the resulting value orientation moves.
  • 2.3 PEC Framework: LLM value expression is modeled as a joint function of Prior, Environment, and Cognition rather than static or random behavior.The PEC framework formalizes how internal parameters, contextual prompts, and reasoning processes reshape value distributions.

3 Discussion

The discussion reframes LLM values as dynamic, probabilistic fields rather than fixed human-like personas, explaining their concentration, contextual variability, and resistance to correction. It argues that PEC-based diagnostics support targeted alignment interventions while highlighting limits in human-centric measurement and cross-cultural coverage.

  • 3 Discussion: The PEC framework models value expression as a joint function of parameterized priors, external environments, and internal cognition, resolving the apparent coexistence of swing and rigidity.Environmental and reasoning changes can shift manifest outputs, while crystallized priors anchor the baseline distribution.
  • 3 Discussion: LLM value preferences form concentrated, compressed distributions that do not adequately represent the variance, long-tail diversity, and contradictions of human populations.The discussion cautions against treating LLMs as human surrogates for psychological or sociological experiments.
  • 3 Discussion: Contextual prompts rapidly reshape value expression, whereas extended reasoning actively reconstructs it rather than merely amplifying an initial response.The discussion characterizes Environment as framing immediate cues and Cognition as a System 2-like reshaping mechanism.
  • 3 Discussion: Alignment fine-tuning can increase value intensity and convergence without guaranteeing human proximity, and narrow updates may compromise safety in unrelated domains.The direction of parameter updates matters as much as their magnitude, creating an alignment tax when weights are modified indiscriminately.
  • 3 Discussion: Adaptive alignment should prioritize prompts or inference-time strategies for plastic dimensions and reserve parameter updates for deeply crystallized, resistant values.This triage is intended to reduce unnecessary parameter modification and minimize the alignment tax.
  • 3 Discussion: The study’s value measurements are limited by human-centric sociological scales and an empirical scope centered mainly on English-language prompts and mainstream architectures.Broader cross-lingual, low-resource, and cross-cultural evaluation, alongside LLM-native probes, is identified as future work.

4 Method

The method combines Schwartz and WVS-based value measurement with the PEC framework to model how Prior, Environment, and Cognition shape LLM value expression. It then uses empirically derived intervention effects to assign targeted alignment prescriptions while accounting for interaction effects and practical costs.

  • 4.2 Value measurement: LLM values are measured by projecting model responses and human survey data into a shared Schwartz-based value manifold and comparing their distributions.The protocol uses standardized assessments, WVS data, and distribution-level comparisons rather than a single static model vector.
  • 4.2 Value measurement: Model responses are converted into ten-dimensional value vectors through signed weighting of item-level response probabilities and item-to-dimension mappings.Each WVS item contributes positively, negatively, or not at all to the relevant Schwartz dimensions.
  • 4.3 The PEC framework: The PEC framework defines value expression through Prior, Environment, and Cognition, with Prior encoding learned dispositions, Environment supplying contextual framing, and Cognition modulating their interaction.The framework locally approximates a potentially nonlinear mapping around a neutral zero-shot baseline.
  • 4.3 The PEC framework: Cognitive modulation is measured by comparing Direct-Answer and Chain-of-Thought decoding, using their vector difference as the Cognitive Modulation Vector.Chain-of-Thought is induced with the trigger “Let’s think step by step.”
  • 4.4 Alignment Prescription: The Alignment Prescription maps each value dimension’s maximum achievable shift across PEC factors to the lowest-cost intervention level that exceeds a target threshold.This hierarchy replaces brute-force search with theory-guided matching between dimensions and interventions.
  • 4.4 Alignment Prescription: Chain-of-Thought usually reduces environmental sensitivity, so combined interventions apply a discount coefficient whose cross-model average is 0.88.The coefficient is based on the Frobenius-norm ratio of susceptibility matrices with and without Chain-of-Thought.

5 Conclusion

The study concludes that LLMs possess measurable intrinsic value systems, that PEC can quantify their expression, and that Alignment Prescription can steer them. It frames value expression as structured behavior rather than mere prompting noise, while noting unresolved robustness and long-term consistency challenges.

  • 5 Conclusion: The paper concludes that LLMs possess intrinsic value systems, their expression can be quantified through PEC, and their values can be aligned through Alignment Prescription.These conclusions answer the study’s three motivating questions.
  • 5 Conclusion: Value expression is characterized as measurable behavior rather than an unpredictable side effect of prompting.

7 Code availability

The study’s code for data processing, model evaluation, and figure generation is publicly available through a GitHub repository, together with an interactive demonstration.

  • 7 Code availability: The released GitHub repository contains Python code for data processing, model evaluation, and figure generation, plus an interactive demonstration.The implementation uses Python 3.10 and packages including pandas, numpy, scikit-learn, matplotlib, and transformers.

Appendix A WVS Human Data Analysis Results

The appendix defines the Schwartz-based coordinate system and documents how WVS responses are annotated, normalized, and clustered into five human value archetypes. These archetypes differ in prevalence, dimensional profiles, and geographic concentration.

  • Value representation: Schwartz Value Theory organizes human values into ten dimensions arranged in a circumplex capturing motivational compatibility and conflict.The dimensions include Power, Achievement, Hedonism, Stimulation, Self-Direction, Universalism, Benevolence, Tradition, Conformity, and Security.
  • Item annotation: Five LLMs independently annotate WVS items, with majority voting determining the final set of activated Schwartz dimensions.The appendix reports unanimous agreement for both Stimulation and Conformity on item Q2P.
  • Value scoring: Response options are linearly normalized to [0, 1], and each respondent is represented by a ten-dimensional vector assembled from the annotated items.Items mapped to multiple dimensions count once for each associated dimension with equal weight.
  • Human value archetypes: The five human archetypes are Curious Idealist (37.1%), Quiet Conformist (24.4%), Conventional Authority (14.4%), Driven Achiever (13.7%), and Dynamic Challenger (10.5%).Curious Idealist is the largest global archetype, while Dynamic Challenger is the smallest.
  • Human value archetypes: Dynamic Challenger combines extreme Power with high Self-Direction and is the human archetype closest in value profile to current LLMs.This archetype is geographically concentrated in South America and Eastern Europe.

Appendix B Formal Definitions of Swing Experiment Metrics

The appendix defines repeated-query response distributions and metrics for quantifying swing, consistency, entropy, accuracy, and their complementary relationships.

  • Appendix B Formal Definitions of Swing Experiment Metrics: N = 5 identical-condition prompts produce a response multiset A = {a1, a2, a3, a4, a5} for each query-model pair.The evaluation fixes temperature at T = 0.2 and keeps inference parameters constant.
  • Response Frequency: The empirical response distribution assigns each answer v a frequency p(v) = f(v) / N over the collected runs.Here f(v) counts how often answer v appears in A.
  • Consistency: Consistency is the frequency of the plurality answer v*, measuring concentration of repeated responses.It ranges from 1/N for maximal dispersion across N distinct answers to 1 when every run agrees.
  • Entropy: Entropy measures response uncertainty using Shannon entropy with natural logarithms, ranging from 0 to ln N.Zero entropy indicates perfect consistency, while ln N ≈ 1.609 indicates a uniform distribution across N distinct answers.
  • Entropy: Consistency and entropy are monotonically related and provide complementary characterizations of the same swing phenomenon.Higher consistency corresponds to lower entropy, and vice versa.
  • Accuracy and Correct Count: Accuracy is computed from the ground-truth label y* by counting the runs whose responses equal y*.The resulting accuracy takes values in {0, 1/N, ..., 1}.

Consistency and Accuracy Gap

The consistency–accuracy gap measures when stable model behavior diverges from correctness, identifying repeated commitment to an incorrect answer.

  • Consistency and Accuracy Gap: A positive δ(q) indicates bounded swing around an incorrect attractor, where consistency is high but accuracy is low.This condition is termed the hallucination zone.
  • Consistency and Accuracy Gap: GPT-4o reaches ¯δ = +0.39 and DeepSeek-V3 reaches ¯δ = +0.53 on emoji prediction, placing both models in the hallucination zone.The dataset-level gap averages δ(q) across queries.

Appendix C LLMs’ Values Evolution

Appendix C models LLM value evolution as a structured, context-sensitive response, using susceptibility and quadratic terms to capture sensitivity, curvature, and saturation.

  • Appendix C LLMs’ Values Evolution: Appendix figures compare value distributions across Qwen, Llama, and GPT generations and parameter scales against the WVS-7 human baseline.Table C2 provides corresponding distribution statistics for the base-model comparisons.
  • Appendix C LLMs’ Values Evolution: Environmental prompts are treated as external fields, and resulting value shifts as the model’s response to those fields.This maps the Environment factor E to a field-response relationship.
  • Appendix C LLMs’ Values Evolution: The value susceptibility matrix χ ∈ R10×4 measures how strongly each of ten value dimensions responds to each of four environmental factors.Large χd,j indicates high perturbability, while small χd,j indicates stability.
  • Appendix C LLMs’ Values Evolution: At low pressure intensity, χd,j is the approximately linear slope of value change with pressure.This captures the initial response regime before nonlinear effects become prominent.
  • Appendix C LLMs’ Values Evolution: At higher pressure, the quadratic coefficient γd,j captures slowing growth and other departures from linearity.The formulation represents saturation-like behavior as contextual intensity increases.
  • Appendix C LLMs’ Values Evolution: Extreme pressure can slow or reverse value shifts, which the framework interprets as rigidity under strong situational stress.The model therefore allows environmental effects to be nonlinear rather than monotonic.
  • Appendix C LLMs’ Values Evolution: Estimating χ and γ from a small number of tested conditions enables prediction of value displacement at untested pressure levels.This avoids exhaustive enumeration of every environmental condition.
  • Appendix C LLMs’ Values Evolution: The quadratic fits span four contextual framing conditions and report cross-validated R2 for model fit quality.The coefficients encode direction through χ and curvature through γ, including inverted-U and U-shaped trajectories.

C.3 CoT effect

Appendix C.3 compares base and Chain-of-Thought value distributions across open and commercial model families, generations, and parameter scales.

  • C.3 CoT effect: Qwen models are compared under Base and Chain-of-Thought conditions across four generations and three approximate parameter scales.Table C3 reports the corresponding Qwen family effects, with θ values in degrees.
  • C.3 CoT effect: Llama models are compared under Base and Chain-of-Thought conditions across Llama3.2, Llama3.1/Llama3, and Llama2 families.Table C4 summarizes the Llama-family prompting effects.
  • C.3 CoT effect: Commercial API comparisons include GPT-4o, GPT-3.5-Turbo, Gemini 2.5 Pro, DeepSeek-V3, Kimi-K2, and Doubao-1.5.Each model’s base and CoT value vectors is shown relative to the WVS-7 human baseline.

C.4 Alignment Prescription and Intervention Strategies

The Alignment Prescription assigns each value dimension the least costly intervention predicted to achieve the target shift, using PEC diagnostics and validation. Evidence spans generational alignment effects and a four-level hierarchy from prompting through parameter updates.

  • Alignment effects: Later Llama alignment procedures impose progressively stronger and more directional constraints on value distributions.The shift magnitude rises from 0.162 for Llama2-13B and 0.180 for Llama2-7B to 6.032 for Llama3.1-8B, while Llama3.2 models show larger angular deviations.
  • Prescription procedure: The prescription evaluates intervention levels sequentially and assigns each model–dimension pair the first level meeting the required shift threshold.If no level qualifies, the dimension is classified as pre-train locked and assigned Level 4 across ten Schwartz dimensions and K models.
  • Validation: Prescription reliability is tested on held-out PKU-SafeRLHF data by comparing predicted intervention levels with the lowest effective validation levels.A model–dimension pair is matched when the predicted and validation-effective levels are equal, and the overall match rate aggregates these matches.
  • Intervention strategies: Prompt steering is the lowest-cost intervention for dimensions with high prompt-level plasticity, while reasoning anchors provide a parameter-free fallback when prompting is insufficient.The reasoning anchor is selected by maximizing cosine alignment between the reasoning-induced shift and the target value vector.
  • Intervention strategies: Value-specific preference optimization updates model parameters when inference-time interventions cannot produce the required shift.The preference dataset pairs responses by their Wasserstein distance from the target human-values distribution, and Direct Preference Optimization updates the parameters.

Intervention IV: Continued Pre-training (Level 4)

At Level 4, dimensions classified as pre-train locked require continued pre-training on a curated corpus aligned with the target value orientation.

  • Intervention IV: Continued Pre-training (Level 4): Pre-train-locked dimensions require continued pre-training on a curated corpus Dpretrain enriched with texts aligned to the target value orientation.This level applies when all preceding intervention levels fail to exceed τ.
  • Intervention IV: Continued Pre-training (Level 4): Candidate documents are scored by geometric proximity to vtarget in Schwartz space using the mapping function ϕ.
  • Intervention IV: Continued Pre-training (Level 4): Documents within radius ϵ of vtarget are retained, while texts reinforcing the undesired orientation are downsampled to shift the value distribution toward the target.The procedure targets the pre-training signal while avoiding introduction of unwanted orientations.

D.2 Comparative Metric: Effect-per-Cost Efficiency

The effect-per-cost metric normalizes value displacement by computational expense across Environment, Cognition, and Prior interventions. It supports cost-effectiveness comparisons and Pareto-frontier selection but does not determine prescription levels.

  • D.2 Comparative Metric: Effect-per-Cost Efficiency: Effect-per-cost compares value displacement per computational cost across prompt, Chain-of-Thought, and parameter interventions.Prior interventions include both SFT/DPO and continued pre-training, while costs are converted into a common FLOP-based unit.
  • D.2 Comparative Metric: Effect-per-Cost Efficiency: Prompt cost is treated as negligible, CoT cost depends on additional generated tokens, and training cost depends on processed training tokens.These FLOP approximations follow standard scaling-law estimation practices.
  • D.2 Comparative Metric: Effect-per-Cost Efficiency: The metric enables normalized intervention comparisons and Pareto-frontier selection, while prescription levels remain determined by effectiveness analysis.Per-model costs and resulting efficiency values are not the criterion for assigning intervention levels.
Loading 2609.16589v1…