Source-linked AI summary

The Shape of Power: A Multilingual Framework for Social Power Reasoning in Dialogues

Farah Atif, Sougata Saha, Monojit Choudhury

arXiv:2608.28144v1cs.AI

TL;DR

Computational studies lack broad, culturally situated resources for analyzing social power in dialogue. The paper introduces a theoretically grounded multilingual framework with native-speaker annotation and scalable tooling, then applies it to French and Egyptian Arabic movie scenes. Humans agree more on observable attributes than interpretive dimensions, while models show persistent gaps in relational and theory-of-mind reasoning.

  • Problem

    Computational studies and conversational corpora remain limited in linguistic, cultural, demographic, and relational coverage for studying social power across societies.

  • Method

    The paper combines a social-science-grounded schema, pilot-refined native-speaker annotation pipeline, custom interface, and multilingual movie-dialogue corpus.

  • Results

    Human agreement is strong for observable demographic and contextual variables but more contested for power asymmetry and intention alignment, while six LLMs and MLLMs show gaps in relational and theory-of-mind reasoning.

  • Takeaways & Limitations

    The framework provides an extensible setting for cross-cultural social-power analysis and evaluation of social reasoning in dialogue.

  • Takeaways & Limitations

    The current dataset covers only French and Egyptian Arabic, restricting generalizability across broader cultural settings.

Abstract

from arXiv · show

Social power plays a fundamental role in shaping human interaction, yet computational studies of power remain limited to narrow linguistic and cultural settings. Existing datasets further lack the demographic and relational depth needed for robust cross-cultural analysis. To address this gap, we introduce a theoretically grounded framework for studying social power in naturalistic multilingual dialogue through movie screenplays. The framework integrates a schema informed by social science theory, a native speaker annotation pipeline refined through pilot studies, and a custom interface for scalable cross-lingual analysis. Using this framework, we constructed an initial corpus containing 15,836 annotated instances from 100 scenes in French and Egyptian Arabic movies. Our analysis reveals strong agreement on observable demographic and contextual attributes, while socially interpretive aspects, such as power asymmetry and intention alignment, remain more contested, highlighting the complexity of social power across cultures. We evaluated 6 Large Language Models (LLMs) and Multimodal LLMs on cross-cultural social power reasoning, finding persistent gaps between human and model agreement in relational and theory-of-mind reasoning. Our work introduces the first extensible multilingual framework for studying social power in dialogues and provides an initial evaluation setting for studying cross-cultural social reasoning.

1 Introduction

Computational research has not systematically modeled social power across cultures, while existing corpora lack the demographic, relational, and cultural depth needed for such analysis. The paper introduces a theoretically grounded, language-agnostic framework and SocLens, a multilingual dataset for studying power in naturalistic dialogue.

  • Existing computational studies commonly treat power as monolithic and tied to a single language, social context, or interaction type.
  • Current conversational corpora lack the demographic, relational, and cultural depth needed to study how power is expressed, perceived, and negotiated across societies.
  • The framework operationalizes social power through theoretically grounded demographic, relational, and theory-of-mind variables, native-speaker annotation, iterative pilots, and a custom cross-lingual interface.
  • SocLens contains 100 densely annotated screenplay scenes in French and Egyptian Arabic, with independent native-speaker annotations and benchmarks against text and multimodal LLMs.

2 Social Power: A Primer

Social power has been theorized as control, the ability to impose one’s will, and influence over values and legitimacy. Computational work has studied linguistic and structural signals, but remains limited in multidimensional, culturally situated modeling.

  • Marxist theory roots power in control over the means of production, while Weber defines it as carrying out one’s will within a social relationship despite resistance.
  • Mann’s ideological power concerns shaping values, beliefs, and perceptions, whereas Bourdieu’s symbolic power operates through naming and classification that confer status and legitimacy.
  • Computational studies have examined power through language mirroring, politeness, network structure, and organizational email hierarchies.
  • The framework addresses gaps by modeling multidimensional power with demographic, relational, contextual, and perspective-dependent features across multilingual, culturally situated interactions.

3 Annotation Framework

The annotation framework uses movie screenplays as naturalistic, culturally situated dialogue and combines social-power theory with human-operationalizable, LLM-assessable, and cross-lingually extensible schema design.

  • Screenplays provide human-crafted dialogue in believable social contexts, standardized formatting, longitudinal character trajectories, and culturally situated interpersonal dynamics.
  • The schema selects features with theoretical or empirical connections to power, including social class, occupation, age, gender, relationship type, and goal alignment.
  • Pilot agreement and annotator feasibility guide schema revision, while precise feature definitions support zero-shot or few-shot LLM assessment.
  • The framework is designed to remain language-agnostic and culturally broad, adding culturally specific alternatives where universal categories are insufficient.
  • Annotations cover speaker demographics, emotional states, power types, interaction dynamics, and dialogue context.

4 Dataset Creation

Dataset creation transforms screenplay text into verified, structured scenes through segmentation, filtering, annotation-tool development, pilot refinement, and large-scale human annotation. The corpus statistics reflect differences in scene composition across the French and Arabic subsets.

  • The pipeline proceeds through movie and script selection, scene segmentation and filtering, feature and guideline design, pilot studies, and large-scale annotation.
  • Scene preparation uses regular expressions for screenplay markers and an LLM for dialogue extraction, scene details, and cumulative summaries.
  • Manual verification removes irrelevant or explicit scenes and corrects speaker attribution, dialogue segmentation, scene mapping, and timestamps.
  • Pilot 1 assessed feature clarity, feasibility, initial agreement, and schema problems using textual inputs, while Pilot 2 added audiovisual cues and revised guidelines.
  • Native-speaker annotators received feature training and interface demonstrations before approximately three weeks of large-scale annotation.
  • Despite fewer French scenes and utterances, the French subset contains more annotated items because its scenes include more interacting characters and directed speaker-dynamics edges.

5 Evaluation

The evaluation compares human annotations with six text-only and multimodal models across annotation tiers, using native-speaker agreement measures and model coverage. It assesses whether models reproduce human judgments and whether visual context improves power identification.

  • Evaluation design: Two evaluation lines measure inter-annotator agreement and compare six models against human annotations across all annotation tiers.The dataset uses parallel annotations from two independent native-speaker annotators.
  • Evaluation design: Gwet’s AC2 measures categorical and ordinal agreement, while mean Jaccard similarity measures multi-label agreement.The agreement analysis uses Gwet’s AC2 because it is more robust than Cohen’s κ under skewed distributions.
  • Model benchmark: The benchmark evaluates Gemini-3.1-Pro, GPT-5.1, Gemma-3-27B, Qwen-2.5-14B, Gemini-3.1-Pro with video, and Molmo2.Text-only evaluation includes two closed-weight and two open-weight models; multimodal evaluation uses video input for two models.
  • Model benchmark: Each model receives the dialogue, scene summary, prior-scene summary, and full annotation schema, with experiments run on an NVIDIA RTX 6000 PRO machine.The estimated cost is approximately 100 USD each for GPT-5.1 and Gemini models.

6 Analysis

Human and model agreement is strongest for observable demographic and contextual attributes and weakest for relational, implicit, and culturally grounded social inference. Coverage limitations and recurring power-type confusions further constrain interpretation of multimodal comparisons.

  • 6.1 Human and LLM Comparison: Gender agreement approaches or reaches 1.00 across most systems, while relationship category, familiarity, and educational tier also show comparatively strong model agreement.Models perform better when social cues are explicit or directly recoverable from dialogue and context.
  • 6.1 Human and LLM Comparison: Human agreement is high for Egyptian Arabic socio-economic class (0.93), social class (0.90), relationship category (0.97), and familiarity (0.83).Agreement is substantially lower for power difference, social-status asymmetry, and intention alignment, which require relational inference and perspective-taking.
  • 6.1 Human and LLM Comparison: The largest human–model gaps involve implicit power structures, perspective-taking, and culturally grounded social inference.In Egyptian Arabic, religion agreement ranges from 0.19 to 0.53 among text models; GPT-5.1 achieves only 0.12–0.15 agreement on French social-status difference.
  • 6.1 Human and LLM Comparison: GPT-5.1 and Gemini-3.1-Pro achieve 100% coverage, whereas Gemini-3.1-Pro with video and Molmo2 cover only 55.7%.Multimodal scores therefore reflect only successfully annotated items and cannot establish a general benefit from multimodal input.
  • 6.2 Which Power Types Are Most Contested?: Referent/Charismatic power is the most contested category, with disagreement often arising when no identifiable power label is assigned.Explicit-label disagreements include Expert versus Referent/Charismatic, Legitimate versus Referent, and Coercive versus Legitimate distinctions.
  • 6.2 Which Power Types Are Most Contested?: Disagreements at the legitimate–coercive boundary reveal conceptual ambiguity rather than necessarily indicating annotation error.In a film centered on an abusive father, one annotator interpreted behavior as paternal authority while another viewed it as domination and coercion.
  • 6.2 Which Power Types Are Most Contested?: Coercive and Reward-based power are most often confused with Legitimate/Legal authority, while Referent/Charismatic and Expert power are frequently reinterpreted as Informational or Expert authority.These recurring confusions are reported across models over the French and Egyptian Arabic data.

7 Discussion and Conclusion

The framework makes social power the primary object of multilingual annotation, enabling cross-cultural and theory-of-mind analysis. Human agreement is stronger for observable cues than interpretive dimensions, while models still struggle with deeper socially grounded reasoning.

  • The framework combines a social-science-grounded schema, native-speaker annotation pipeline, and custom interface for cross-lingual extensibility.
  • Annotating social power directly captures relational features from both annotator and character perspectives.
  • Human annotators agree strongly on observable demographic and contextual variables but contest power asymmetry and intention alignment more often.
  • Current LLMs handle explicit social cues reasonably well but struggle with cultural understanding and theory-of-mind inference.

8 Limitations

The study’s findings are bounded by its limited linguistic coverage, small annotator pool, rapidly changing model landscape, and multimodal evaluation failures.

  • The dataset covers only French and Egyptian Arabic, limiting generalizability across broader cultural settings.The authors do not claim that power types are associated with these cultures.
  • Two annotators per scene cannot fully capture plausible interpretations of contested dimensions such as power asymmetry and intention alignment.A larger annotator pool would support deeper analysis of perceptual variation and demographic correlates.
  • The six-model benchmark may date quickly because the LLM and multimodal landscape evolves rapidly.The authors plan to support replication by releasing the schema, prompts, and annotated data.
  • Safety-aligned refusals and structural output failures reduce effective coverage in some multimodal evaluations.These failures should be considered when interpreting agreement scores.

9 Ethical Considerations

The work addresses sensitive demographic attributes and interpersonal power relations using trained, culturally familiar annotators and guidelines against unsupported assumptions. However, media representations and societal norms may still introduce cultural bias.

  • The framework annotates sensitive attributes and interpersonal relations, including social class, religion, and hierarchy.
  • Trained native speakers familiar with relevant cultural contexts performed annotations under guidelines discouraging unsupported assumptions.
  • Human annotators and LLMs may still reproduce cultural biases present in media representations and societal norms.
  • Explicit or harmful scenes were filtered where appropriate, and annotators were informed that some scenes could be sensitive or emotionally difficult.

A Appendix

The appendix supplies supplementary documentation for the framework, including its schema, annotation procedures, and prompts.

  • The appendix details feature definitions spanning demographics, emotional states, power types, speaker dynamics, and dialogue context.
  • It describes annotator recruitment, training, and qualification procedures used during dataset construction.
  • It provides the prompts used throughout the framework.

A.1 Detailed Annotation Scheme

The framework models social power in dialogue as a multidimensional phenomenon spanning observable, relational, contextual, and socially interpretive features. Figure 4 presents its five interconnected components.

  • Social power is modeled through demographic cues, interpersonal relations, contextual grounding, and socially situated reasoning.
  • The schema combines observable attributes such as age and occupation with interpretive variables including power asymmetry and intention alignment.
  • Figure 4 organizes the framework into five interconnected components covering demographic features, emotions, power types, speaker dynamics, and dialogue context.

A.1.1 Demographic features

The annotation scheme records demographic, emotional, power, relational, and contextual properties of dialogue scenes. It combines culturally sensitive observable categories with interpretive judgments from annotators and characters’ perspectives.

  • Demographic features: Demographic features include age group, gender, country of origin, social class, socio-economic class, occupation, education, marital status, religion, and ethnicity.Several categories use broad or coarse-grained groupings to accommodate cultural variation and inference limits in movie scenes.
  • Demographic features: Occupation and educational tier capture visible institutional, expertise, and socio-economic positioning using WVS- and ISCED-informed categories.Education is grouped into Higher, Medium, and Lower tiers because fine-grained attainment is difficult to infer from movies.
  • Emotions: Emotions are annotated at scene level using six primary categories: Love, Joy, Surprise, Anger, Sadness, and Fear.The scheme treats affective information as related to persuasion, conflict escalation, social dominance, and interpersonal alignment.
  • Power types: Power is represented through coercive, reward-based, legitimate, expert, referent, informational, and ideological forms that may coexist and vary across cultures.The framework does not reduce power to institutional authority alone.
  • Speaker dynamics: Relational annotations cover directional power difference, social-status difference, familiarity, goal alignment, and relationship category and type.Power and status judgments can reflect both annotators’ interpretations and characters’ perspectives, supporting analysis of theory-of-mind reasoning, misunderstanding, and deception.
  • Dialogue context: Contextual properties distinguish professional versus personal domains and public versus private interaction spaces.These variables capture how setting, audience presence, and institutional environment shape social behavior and power expression.
Loading 2608.28144v1…