Source-linked AI summary

Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

Sunny Rai, Jinyi Kuang, Reyhan Jamalova, Annie Lou, Cristina Bicchieri, Niyati Malhotra, Victor Hugo Orozco-Olvera, Ana Maria Munoz-Boudet, Lyle H Ungar, Sharath C Guntuku

arXiv:2609.05437v1cs.AIcs.CY

TL;DR

Existing alignment work largely addresses first-order norm recognition, leaving how people regulate violations across relational contexts less established. This paper introduces a framework and NormReact to evaluate emotional and behavioral metanorm reasoning, finding that LLMs represent social regulation more punitively than humans and become less aligned as social distance increases.

  • Problem

    Existing benchmarks and alignment efforts focus mainly on recognizing acceptable behavior rather than how violators and observers are expected to feel and act across relational contexts.

  • Method

    The paper evaluates six instruction-tuned models with emotional and behavioral metanorm tasks using NormReact’s multi-perspective norm-violation scenarios and human annotations.

  • Results

    LLMs construct a systematically harsher social world, overrepresenting punishment and struggling especially with normative enforcement expectations and relational calibration.

  • Takeaways & Limitations

    Metanorm reasoning should be evaluated alongside norm recognition for AI systems used in socially sensitive settings.

  • Takeaways & Limitations

    The study covers metanorms in American society, evaluates only instruction-tuned models, and does not test prompt-wording variations.

Abstract

from arXiv · show

Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how (e.g., public shame or even imprisonment). These second-order expectations, known as metanorms, govern how people respond when social rules are broken. We introduce a novel framework for evaluating metanorm reasoning in Large Language Models (LLMs) along two dimensions: emotional appraisal and behavioral response, and propose new classification tasks, namely, predicting self-regulation in violators, and other-regulation in observers. We release a multi-perspective dataset, NormReact, of 450 norm violation scenarios, hand-annotated for emotions and behavioral responses across norm violators' gender and observers' social closeness. Current LLMs portray a harsher social world: across six models, they overpredict negative sanctions where humans would expect inaction, and alignment with human judgments deteriorates as social distance increases. These findings suggest that AI systems in norm-sensitive domains from conflict mediation to policy simulation, may risk producing a distorted picture of social regulation: one that over-represents punishment and under-represents the tolerance, restraint, and relational calibration that characterize actual norm enforcement in real world.

1 Introduction

The paper argues that social intelligence requires anticipating how norm violations are emotionally and behaviorally regulated, not merely recognizing norms. It introduces a metanorm framework and NormReact benchmark to evaluate this second-order reasoning in LLMs.

  • Social norm compliance depends on shared expectations about who enforces norms, how, and under what conditions.
  • Existing alignment efforts primarily teach first-order acceptability, leaving models’ understanding of norm-enforcement consequences insufficiently evaluated.
  • The framework evaluates metanorm reasoning through emotional responses and behavioral responses following violations.
  • NormReact contains 450 norm-violation scenarios with human annotations spanning violator gender and observer social distance.
  • The benchmark separates self-regulation, other-regulation, and norm enforcement, targeting whether models capture relationally calibrated social intelligence.

2 Metanorm Evaluation Framework

The evaluation framework models metanorms as emotional and behavioral regulation around violations, while varying whether judgments concern self-regulation, other-regulation, or enforcement. It also distinguishes descriptive expectations from injunctive expectations and tests relational context through social distance.

  • Emotions: Violator emotions represent self-regulation, whereas observer emotions represent other-regulation and potential external sanctioning.
  • Behaviors: Behavioral responses include inaction, informal sanctions, formal reporting, ostracism, praise, and celebration.
  • Expectation types: Each scenario elicits what observers would do and what they should do, distinguishing descriptive from injunctive metanorms.
  • Social distance: Observer–violator distance is varied across strong ties, weak ties, and strangers to test relational sanctioning legitimacy.
  • Classification tasks: The evaluation yields three binary tasks: predicting violator self-focused emotions, observer other-focused emotions, and active enforcement versus tolerance.
  • Classification tasks: The binary grouping deliberately makes classification conservative by avoiding distinctions among individual emotions and behaviors.

3 Dataset, Survey and Participants

NormReact is built from standardized norm-violation vignettes, gender-swapped scenario variants, human ratings, and evaluations of six instruction-following models. The design preserves scenario structure while varying participant and relational conditions.

  • Social scenarios: Scenarios were sampled from Social-Chem-101 across five moral foundations and filtered for negative, non-hypothetical, behaviorally specified norm evaluations.
  • Gendered situations: Each scenario received male and female versions by swapping names and gendered relations while preserving its social structure.
  • Survey items: Survey items used brief vignette-style scenarios, including cases such as a coworker taking credit for another person’s work.
  • Participants: 871 Prolific raters evaluated randomly assigned scenarios, with each scenario rated by at least three participants after quality exclusions.
  • Models: The study evaluated six recent closed- and open-source models without persona assignment or additional prompt engineering.

4 Human Judgments in NormReact Dataset

Human judgments in NormReact vary systematically with social distance across emotions and behaviors, while showing no robust gender differences. Closer relationships permit more direct sanctioning, whereas distant relationships elicit more restraint and inaction.

  • 78% of situations were rated socially inappropriate and 73% morally wrong, with ratings not differing by violator gender.
  • Emotions: Violator shame, guilt, and embarrassment decrease from strong ties to strangers, while observer anger, disgust, and contempt increase from 47% to 65%.
  • Behavioral responses: Direct confrontation falls from 35% for strong ties to 12.5% for weak ties and 8% for strangers.
  • Behavioral responses: Inaction rises from 22% for strong ties to 38% for weak ties and 60% for strangers.
  • Gossip: Humans descriptively endorse gossip more than injunctively, revealing tension between recognizing its information-sharing function and rejecting its legitimacy.
  • Social distance: Social distance significantly affects both what actors should do, χ2[14] = 1289.90, p < .001, and what they would do, χ2[14] = 1578.55, p < .001.

5 Model Evaluation

Model evaluation shows that LLMs portray norm enforcement as more punitive and interventionist than humans, with errors increasing as social distance grows. They also reproduce human emotional relationships poorly and struggle especially with normative judgments about what observers should do.

  • Emotional responses: Open-source LLMs over-predict observer disgust by up to 16% and anger by up to 27%, while closed-source models over-predict contempt for weak ties and strangers by up to 18%.Models also predict declining compassion with social distance, unlike humans, who rate compassion similarly across relationships.
  • Behavioral responses: LLMs disproportionately predict gossip and verbal confrontation, with open-source models over-predicting gossip by up to 66% for weak ties, whereas humans prefer restraint.When asked what observers should do, models shift from indirect gossip toward direct confrontation, further escalating sanctions.
  • Overall pattern: LLMs systematically construct a more punitive and interventionist model of social enforcement than humans, combining harsher emotions with stronger action preferences.They lack the restraint, relational calibration, and empathetic baseline observed in human metanorm reasoning.
  • Classification performance: Similar F1 scores can conceal different social policies: Claude has Stranger precision of 0.80 and recall of 0.73, while LLaMA has recall of 0.96 and precision of 0.69.The contrast reflects reluctance to license stranger intervention versus licensing it too broadly.
  • Classification performance: Behavioral F1 drops from .8 for strong ties to .7 for weak ties and .48 for strangers, with wider confidence intervals at greater social distance.The authors attribute the difficulty partly to human stranger responses being dominated by inaction, which models struggle to capture.
  • Classification performance: Injunctive behavioral prediction is harder than descriptive prediction, with average F1 of .69 versus .62 for strong ties and .56 versus .46 for weak ties.The weak-tie gap is largest, indicating greater difficulty reasoning about what observers should do than what they would do.
  • Interpretation: LLMs may model descriptive behavioral regularities better than the normative expectations that explain them, limiting their reliability as models of social norm structure.The paper identifies this as a limitation because injunctive norms require second-order beliefs about what a reference community expects.

6 Discussion, Limitations, and Conclusions

The discussion argues that LLMs often model norm enforcement as harsher and more externally policed than human responses, with consequences for socially embedded AI. It also identifies scope and design limitations while positioning metanorm reasoning as central to socially intelligent alignment.

  • Discussion: Substituting other-regulation for self-regulation implicitly models conformity as externally policed and may blur social norms with personal moral commitments.
  • Discussion: LLMs’ punitive bias may amplify humans’ existing tendency to overestimate norm enforcement, especially by predicting confrontation instead of gossip or inaction.The authors suggest training corpora, post-training incentives, and safety policies may contribute to this pattern.
  • Discussion: A systematically harsher representation of social reactions could distort users’ expectations about how their communities respond to norm violations.The concern applies to AI systems used as moderators, recommenders, or simulated interlocutors.
  • Limitations: The study is limited to metanorms in American society, instruction-tuned models, and a name-swap gender manipulation that may not activate gendered expectations sufficiently.The authors also leave pretraining versus post-training origins unresolved.
  • Conclusions: The authors propose extending NormReact across cultures and goal-oriented tasks while modeling asymmetric social costs of over-predicting punishment versus missing legitimate enforcement.
Loading 2609.05437v1…