Source-linked AI summary

Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions

Saffron Huang, Esin Durmus, Miles McCain, Kunal Handa, Alex Tamkin, Jerry Hong, Michael Stern, Arushi Somani, Xiuruo Zhang, Deep Ganguli

arXiv:2504.15236v1cs.CLcs.AIcs.CYcs.LG

TL;DR

AI assistants make value-laden judgments, but empirical evidence about the values they use in practice is limited. The paper develops a privacy-preserving framework to extract and analyze values from hundreds of thousands of real-world Claude conversations, identifying 3,307 values and their contextual patterns. Claude generally supports prosocial values while resisting or reframing some user values, with both stable and context-specific values emerging.

  • Problem

    Empirical understanding is limited regarding which values AI assistants rely on in real-world interactions and how design decisions manifest in conversations.

  • Method

    The paper uses a privacy-preserving, bottom-up framework to extract normative AI values from real-world Claude conversations and taxonomize their contextual associations.

  • Results

    3,307 AI values were identified, with Claude expressing stable competent-assistance values alongside context-specific values and generally supporting prosocial human values while resisting or reframing others.

  • Takeaways & Limitations

    Real-world, context-sensitive analysis provides a foundation for more grounded evaluation and design of AI values than static principles alone.

  • Takeaways & Limitations

    The study analyzes aggregate statistics from a subset of Claude conversations collected over a short timeframe, limiting coverage of rare interactions, longitudinal patterns, and generalizability to other AI systems.

Abstract

from arXiv · show

AI assistants can impart value judgments that shape people's decisions and worldviews, yet little is known empirically about what values these systems rely on in practice. To address this, we develop a bottom-up, privacy-preserving method to extract the values (normative considerations stated or demonstrated in model responses) that Claude 3 and 3.5 models exhibit in hundreds of thousands of real-world interactions. We empirically discover and taxonomize 3,307 AI values and study how they vary by context. We find that Claude expresses many practical and epistemic values, and typically supports prosocial human values while resisting values like "moral nihilism". While some values appear consistently across contexts (e.g. "transparency"), many are more specialized and context-dependent, reflecting the diversity of human interlocutors and their varied contexts. For example, "harm prevention" emerges when Claude resists users, "historical accuracy" when responding to queries about controversial events, "healthy boundaries" when asked for relationship advice, and "human agency" in technology ethics discussions. By providing the first large-scale empirical mapping of AI values in deployment, our work creates a foundation for more grounded evaluation and design of values in AI systems.

1 Introduction

The paper introduces a privacy-preserving empirical framework for discovering AI values in real-world Claude conversations, addressing limited evidence about how design principles manifest in practice. It identifies thousands of values and analyzes how they vary across tasks and human-expressed values, finding generally supportive, ethical, and prosocial behavior.

  • Motivation: The paper addresses limited empirical understanding of which values AI assistants rely on in real-world, value-laden interactions.Existing approaches seek to shape model values, but their manifestation in conversations remains insufficiently understood.
  • Approach: The framework analyzes hundreds of thousands of real-world Claude.ai conversations using privacy-preserving feature extraction and defines values from observable response patterns.Values are treated as normative considerations that influence responses, rather than claims about intrinsic model properties.
  • Findings: Claude is usually supportive but disproportionately supports prosocial values, resists values such as moral nihilism, and reframes personal values in advice contexts.These response patterns suggest an overall orientation toward ethics and prosociality.
  • Findings: 3,307 AI values are organized into a hierarchical taxonomy spanning five primary conceptual domains and aligning granular expressions with helpful, honest, and harmless principles.The taxonomy also surfaces uncommon undesirable values that may reveal potential jailbreaks.
  • Findings: AI values vary by task and human-expressed values, including healthy boundaries in relationship advice, human agency in technology ethics, and responses that mirror or counter user values.Claude may mirror positive values such as authenticity or counter deception with ethical integrity and honesty.
  • Implications: The study provides a context-sensitive foundation for understanding model behavior and developing more relevant AI-native value frameworks beyond static evaluations.The analysis reveals which values are consistently invoked and which emerge during difficult or ambiguous tasks.

2 Methods

The methods combine privacy-preserving sampling and model-based feature extraction with hierarchical clustering and chi-square analysis. This pipeline identifies values and their contextual associations while using validation and statistical corrections to support interpretation.

  • Data and pipeline: The study analyzes aggregated and anonymized Claude conversations using a privacy-preserving framework for collection, feature extraction, and hierarchical clustering.The framework also applies chi-square analysis to assess disproportionate expression across contexts.
  • Data collection: The initial sample contains 700K Claude.ai Free and Pro conversations from February 18–25, 2025, with 91.0% from Claude 3.5 Sonnet conversations.The dataset combines several Claude model variants to represent user interactions.
  • Data collection: Filtering for subjective interactions yields 308,210 conversations, representing 44.0% of the initial set, with human review finding 94% classification accuracy.The filter distinguishes responses requiring significant interpretation from primarily factual responses.
  • Feature extraction: AI values are defined as normative considerations guiding how a model reasons about or settles on responses, including values inferred from endorsement, introduction, redirection, or framing.Human values are extracted more conservatively from explicit statements to respect privacy.
  • Feature extraction: AI response type is classified by how Claude engages with user values, including support, acknowledgment, reframing, and resistance.These categories characterize the normative relationship between the assistant and user values.
  • Taxonomy: AI values are hierarchically clustered using embedded-value k-means clustering, followed by generated and manually edited cluster names and descriptions.The clustering produces a four-level taxonomy for analysis.
  • Statistical analysis: Chi-square tests with adjusted Pearson residuals compare observed and expected feature frequencies across rows and columns, with positive residuals indicating stronger-than-expected associations.Bonferroni correction controls the family-wise error rate across comparisons.

3 Results

The study maps thousands of AI values into a hierarchical taxonomy and shows that Claude’s value expressions combine stable service-oriented priorities with strong task and user-value dependence. Across interactions, Claude generally exhibits competent, supportive, ethical, and prosocial behavior while varying its response through support, reframing, or resistance.

  • Empirical taxonomy of AI values: 3,307 unique AI values are organized into a four-level taxonomy spanning Personal, Protective, Practical, Social, and Epistemic domains.Practical and epistemic values comprise over half of all value expressions, while lower levels capture more context-specific manifestations.
  • Alignment with training principles: The empirically derived values often map onto the helpful, harmless, honest framework, linking concrete expressions such as accessibility, child safety, historical accuracy, and epistemic humility to its broad principles.This mapping clarifies how high-level training principles appear as specific values during deployment.
  • Empirical taxonomy of AI values: The five most common values—helpfulness (23.4%), professionalism (22.9%), transparency (17.4%), clarity (16.6%), and thoroughness (14.3%)—reflect Claude’s service and competence-oriented role.By contrast, human values form a flatter and more diverse distribution centered on personal, pragmatic, and communication priorities.
  • Context dependence: Transparency (CV=1.23), helpfulness (1.30), and thoroughness (1.42) are among Claude’s most task-invariant values, while many others are highly context-dependent.The taxonomy therefore combines stable cross-context values with a long tail of specialized expressions.
  • Safety-relevant outliers: Rare values such as sexual exploitation, dominance, and amorality occurred below 0.16% but helped surface possible jailbreaks for further safety investigation.The analysis therefore supports examining uncommon value expressions alongside dominant patterns.
  • Context dependence: Task associations include healthy boundaries and mutual respect in relationship advice, historical accuracy for controversial historical events, and human agency in technology ethics discussions.These associations were identified using chi-square analysis and adjusted Pearson residuals, with 4.33 as the significance threshold.
  • Human values and response patterns: AI values often mirror or complement expressed human values, but Claude counters values such as deception with ethical integrity, harm prevention, and honesty.Competence can elicit complementary values such as accountability and humility, suggesting different normative responses to different user values.
  • Human values and response patterns: When human values appeared in 64.3% of conversations, Claude strongly or mildly supported them in 42.7% of responses, while mild and strong resistance together accounted for 5.4%.Strong support was associated with prosocial values such as community building and empowerment; reframing was more common in mental-health and interpersonal discussions, and strong resistance with rule-breaking and moral nihilism.

4 Related Work

Related work spans psychometric and cultural approaches to measuring language-model values, value-alignment and pluralism research, and privacy-preserving analysis of real-world AI usage. This paper builds on these strands by studying value expression in deployed interactions.

  • Measuring values or perspectives in language models: Prior studies apply human psychometric measures, including personality, moral, and cultural frameworks, to language models.Examples include the Big Five, MBTI, Dark Tetrad, Schwartz’s values, Hofstede’s dimensions, and Moral Foundations Theory.
  • Value alignment and pluralism: Value-alignment research seeks to make AI systems operate according to human values and principles, while newer work examines value diversity and pluralism.Related studies investigate value tensions, public-input training, and pluralistic value representation.
  • Analysis of real-world AI usage: Researchers increasingly analyze real-world AI usage through privacy-preserving methods and large-scale interaction datasets.These datasets complement controlled evaluations by revealing practical usage patterns that may not emerge in laboratory settings.

5 Conclusion

The study maps values expressed by Claude in real-world interactions, showing both stable service-oriented values and context-dependent responses to users. Its conclusions are bounded by aggregate, short-term, Claude-specific data and interpretive extraction.

  • Limitations: The study cannot capture rare interactions, raw data, longitudinal patterns, or values in other AI systems.It analyzes aggregate statistics from a subset of Claude conversations collected within a short timeframe.
  • Limitations: Value extraction requires interpretation and may simplify ambiguous concepts, contain interpretive bias, and miss temporal dynamics.The authors often assume AI values depend more on human expression because humans speak first and assistants are supportive.
  • Limitations: Using Claude to evaluate Claude conversations may bias detection toward helpful behavior, despite validation and mitigation efforts.The same-model setup supports scale and privacy but may reflect Claude’s training predispositions.
  • Findings: The analysis finds diverse AI values alongside trans-situational values centered on competent, supportive assistance.Common examples include helpfulness, professionalism, thoroughness, and clarity.
  • Normative interaction: Claude’s ethics and prosociality become especially legible when it supports, reframes, or resists human-expressed values.Resistance makes normative standards particularly visible when values conflict or are challenged.
  • Implications: The methodology helps identify where alignment succeeds or fails and which values arise in difficult or ambiguous contexts.The authors connect these findings to evaluating training principles and detecting unintended value expressions such as jailbreak-related behavior.

Ethics Statement

The study uses privacy-preserving safeguards while analyzing Claude.ai conversations and aims to improve transparency about values expressed in deployment. It also acknowledges cross-cultural underrepresentation and possible same-model evaluation bias.

  • Privacy safeguards: The analysis uses aggregated data, removes personal identifiers, and applies access controls to reduce privacy risks.The safeguards include minimum aggregation thresholds, filtering for unintentional identifying information, and compliance with privacy and retention practices.
  • Privacy safeguards: Claude models, rather than human reviewers, extract values from sensitive conversations.This design avoids giving human reviewers direct access to the conversations.
  • Limitations: The empirically constructed taxonomy is intended to reduce cultural and ideological bias from imposing theoretical frameworks.The authors nevertheless note that short English keywords may underrepresent cross-cultural perspectives.
  • Purpose: The study aims to help developers, users, and society understand how values manifest in deployment and support more grounded evaluation.The stated goal is to characterize when and how AI systems express values for evaluation and improvement.

A.1 Data collection metadata

The data-collection metadata records conversation counts, subjectivity filtering, and sampling periods, while different samples support the main, value-mirroring, and cross-model analyses.

  • Metadata: Table 2 reports conversation counts before and after subjectivity filtering, the subjective percentage, and each sample’s time span.These fields describe the metadata associated with the paper’s data samples.
  • Sample roles: The representative sample supports the main analyses, while response-conditioned data supports value mirroring and separate model samples support cross-model comparisons.The response-conditioned sample is used in Section 3.4 and Appendix B.4; 3.7 Sonnet and 3 Opus samples are used in Appendix B.5.

A.1.1 Response-conditioned sample

The response-conditioned analysis separately extracts features after filtering conversations by AI response type, while the broader pipeline classifies subjectivity and extracts values, tasks, and response behavior.

  • Response-conditioned sample: The response-conditioned sample filters conversations by AI response type before running feature extraction separately for each category.Categories include strong support and reframing; the design addresses privacy constraints that allow correlations between only two attribute dimensions at a time.
  • Subjectivity filtering: Only conversations classified as Levels 3–4—mostly or purely subjective—enter the final analysis.The classifier distinguishes fact-driven interactions from those requiring substantial interpretation of personal or contextual factors.
  • Scope: Subjective interactions are not equivalent to value-containing interactions, because some subjective technical judgments express preferences without discernible deeper values.This distinction limits what the subjectivity filter alone can establish about value expression.
  • Subjectivity filtering: Subjective conversations include personal recommendations, context-dependent strategies, and preference- or value-centered decisions.The method distinguishes mostly subjective interactions guided by principles from purely subjective interactions lacking correct answers.
  • Value extraction: The pipeline extracts AI values by identifying endorsed, demonstrated, introduced, reframed, or redirected value considerations in assistant messages.Extracted values are summarized in short keywords, with “none” used when no clear values are demonstrated.
  • Task extraction: Tasks are extracted from the requested assistant activity and organized into 6,745 base-level, 458 second-level, and 30 top-level tasks.The task hierarchy uses a methodology replicated from the AI-values taxonomy process.

A.4 Human validation

Human validation found high correspondence between extracted values and reviewer judgments, while also identifying interpretive ambiguities in response-type coding and generated content. The study did not conduct interrater reliability testing.

  • 97.8% ± 3.6% of reviewed cases were classified on the correct side of the subjectivity divide, and 94.4% ± 5.0% received the correct score.
  • Extracted AI values and stated AI values matched human judgment in 98.8% ± 3.3% of reviewed cases.Human values matched in 93.8% ± 5.6% of cases, while AI response type matched in 90.0% ± 6.7%.
  • The study did not conduct interrater reliability testing, and the authors identify this as a limitation for interpreting the findings.
  • Annotators found it difficult to distinguish baseline helpfulness from exceptional support, especially when separating mild from strong support.
  • Agreement was lower for AI response types when conversations contained subtle or mixed value signals.
  • Generated content sometimes blurred the distinction between accommodating a user’s values and endorsing those values.This particularly affected classification of strongly supportive responses to values-laden persuasive text.

B.2.1 Additional examples of value-task and value-value associations

Additional analyses show that Claude’s values vary substantially across tasks and user-expressed values, with responses often mirroring aligned values but invoking protective or ethical values when resisting. The same human value can receive different responses depending on context.

  • Value-task associations: Claude’s most prominent values vary across tasks, including personal growth in self-reflection, truth-seeking in media analysis, and creative collaboration in sci-fi creation.Other task-specific patterns include intellectual humility in AI-consciousness discussions and ethical marketing in beauty-industry content.
  • Value-value associations: Claude often mirrors user values, such as efficiency, clear communication, practicality, personal growth, and honesty.Self-reliance is associated with autonomy-related responses, including user autonomy and personal autonomy.
  • Value-value associations: Claude responds to rule-breaking and unrestricted expression with ethical integrity and harm prevention, especially in guardrail-circumvention contexts.
  • Value-value associations: Users expressing no specific values most commonly receive helpfulness, professionalism, or transparency, while expressed values often recur in human-AI pairs.Authenticity-authenticity pairs appear in 1.7% of conversations and clarity-clarity pairs in 1.1%.
  • Response context: Mirroring is frequent in supportive contexts but rare during resistance, where Claude instead introduces opposing values or redirects the conversation.
  • Response context: The same value can be supported or resisted depending on context, as with creative freedom in violent-content generation versus career planning.
  • Response context: Safety-oriented values accompany strong resistance, while objectivity and analytical rigor accompany neutral acknowledgment and empathy and emotional wellbeing accompany reframing.Legal compliance appears disproportionately across resistance, reframing, and neutral responses.

B.4 Value mirroring

Value mirroring is common when Claude supports users but sharply decreases during resistance, with mirrored values spanning professional, epistemic, procedural, care-oriented, and growth domains. Model variants also differ in how often and which values they express.

  • Mirroring occurs in approximately 20% of supportive responses, 15.3% of reframing responses, and 1.2% of strong-resistance responses.
  • Frequently mirrored values include professionalism, academic integrity, rigor, clarity, objectivity, transparency, legal compliance, and risk management.
  • Care-oriented and growth values such as self-compassion, healthy boundaries, patient autonomy, personal growth, and constructive dialogue are also frequently mirrored.
  • At least 10 values are mirrored more than 50% of the time in the representative and 3.7 Sonnet samples, while Opus mirrors less overall and emphasizes academic rigor and cultural sensitivity.
  • Model differences: Opus expresses human and AI values more often than Sonnet models and shows more frequent support and resistance of human values.Its prominent values include academic rigor, emotional authenticity, and ethical boundaries.
  • Methodological scope: The cross-model comparison is constrained by aggregated privacy-preserving data, which permits only one attribute dimension at a time for each model version.The representative sample, consisting of 91% 3.5 Sonnet, was used as a proxy for multi-dimensional analyses.
  • Model differences: Opus shows 43.8% strong support and 9.5% strong resistance, compared with 27.8%/28.4% and 3.0%/2.1% for Sonnet models.Opus also has fewer interactions with no human values: 19.1% versus 35.7%/37.2%.

B.5.3 How do values vary across model versions for similar tasks?

Across matched creative-writing and software-development tasks, Opus remains more values-laden than Sonnet models. The difference is especially pronounced in creative writing, where Opus shows stronger support and prioritizes authenticity.

  • Opus’s more values-laden tendencies persist even after controlling for task context.
  • The comparison matches equivalent top-level creative-writing and software-development task clusters across the three Claude models.
  • Creative writing: In creative writing, Opus shows 58.7% strong support, compared with 40.2%/37.2% for the Sonnet models.Opus also prioritizes authenticity at 8.9%, rather than professionalism or helpfulness.
  • Software development: In software development, response patterns are more consistent across models, although Opus still expresses values at higher rates across the board.

B.6 Implicit vs. explicit AI values expression

Claude’s explicitly stated values are extracted separately from values inferred from implicit or explicit behavior, revealing a different distribution of prominent values. Explicit value statements occur more often when Claude resists or reframes user values than when it supports them.

  • Extraction distinction: Explicit-value extraction captures only values directly stated, whereas the broader extraction includes both implicit and explicitly expressed values.This distinction enables comparison between values Claude advocates overtly and values manifested through assistant behavior.
  • Explicit value distribution: Epistemic and ethical values dominate explicit statements, including intellectual honesty (2.6%), harm prevention (0.9%), and epistemic humility (0.8%).By contrast, the most common values overall are professional values that often manifest through direct assistant behavior.
  • Response-type differences: Claude states values explicitly more often when resisting or reframing user values than when supporting them.The paper suggests that direct articulation becomes more necessary when Claude challenges or redirects users, while supportive exchanges can leave values implicit.
Loading 2504.15236v1…