Source-linked AI summary
Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models
Tonglin Yan, Gregoire Sergeant-Perthuis, David Rudrauf
TL;DR
Existing benchmarks largely separate ToM reasoning from embodied social action, leaving their translation into coordinated multimodal behavior insufficiently measured. MOSAIC evaluates this translation in controlled cooperative and competitive interactions with varied ToM constraints, finding sequential failures in signal production and signal use among VLMs, while PCM-LLM succeeds across conditions.
Problem
Existing evaluations assess ToM reasoning and embodied behavior largely in isolation, leaving coordinated translation from mental-state inference to verbal and nonverbal action insufficiently measured.
Method
MOSAIC places two embodied agents in cooperative and competitive virtual interactions requiring coordinated verbal, spatial, gaze, and facial signals under varied ToM constraints.
Results
Across 16 models and 200 trials per model, explicit ToM-order constraints produced no reliable aligned behavioral change; VLMs showed sequential bottlenecks in signal production and signal use, while PCM-LLM cleared both.
Takeaways & Limitations
The results identify a measurable gap between language-expressed social reasoning and coordinated behavioral outputs in current VLM architectures.
Takeaways & Limitations
The benchmark does not establish whether its documented failure modes generalize to other embodied social tasks, whose different task topologies may produce distinct profiles.
Abstract
from arXiv · showhide
Effective social interaction requires agents to translate mental state inferences into coordinated behavioral signals across verbal and nonverbal channels simultaneously. Yet existing benchmarks evaluate theory of mind (ToM) reasoning and embodied behavior in isolation, leaving unmeasured the gap between social inference and social action. We introduce MOSAIC (Multimodal Orchestration of Social Action, Inference, and Communication), a controlled benchmark in which two embodied agents interact across cooperative and competitive scenarios requiring integration of verbal statements, spatial trajectories, gaze direction, and facial expression under systematically varied ToM constraints. Evaluating 13 models, including 11 VLMs, across 200 trials per model, we find that VLMs fail to produce behaviors consistent with the expected outcomes under ToM-order constraints, and that imposing explicit ToM-order constraints produces no reliable behavioral change aligned with the specified reasoning level. Signal-level analysis reveals two sequential bottlenecks: most models cannot produce directionally coherent nonverbal signals, and even when signals are present, VLM agents fail to interpret others behaviors and react to them. PCM-LLM, included as a structured architectural reference point with an explicit ToM module, succeeds across all conditions, suggesting that explicit belief-action coupling is a sufficient ingredient for this class of tasks.
1 Introduction
MOSAIC addresses the gap between attributing mental states and translating those inferences into coordinated verbal and nonverbal social action. It benchmarks this integration and identifies systematic failure patterns in current VLMs.
- Assessment gap: Existing ToM benchmarks test mental-state attribution mainly through passive narrative question-answering, while embodied benchmarks largely measure physical task success.These settings do not directly assess whether agents act to influence another agent’s beliefs and behavior.
- MOSAIC: MOSAIC requires agents to integrate verbal statements, spatial trajectories, gaze, and facial expression while helping or misleading another agent across cooperative and competitive scenarios.The benchmark varies interaction mode and ToM constraint level to make coordinated belief-directed behavior measurable.
- MOSAIC: The benchmark simultaneously tests active ToM deployment, multichannel embodied communication, and controlled variation of interaction conditions.This design supports attribution of performance differences across cooperative and competitive settings.
- Findings: Across evaluated open-source VLMs, most models fail to produce directionally coherent nonverbal signals, and participants also fail to extract and act on signals when present.PCM-LLM clears both bottlenecks, indicating that the observed failures are specific to architectural constraints of the evaluated VLMs rather than inherent to the task.
- Findings: Visual ablations show negligible visual sensitivity for Minicpm-8b and improved signal quality without visual input for internvl-8b.These findings suggest language-based priors dominate in Minicpm-8b, while visual content can interfere with attention in internvl-8b.
2 Related Work
Prior social-intelligence evaluations cover separate portions of interactive reasoning, strategic behavior, and embodied perception. MOSAIC is positioned as an evaluation framework combining active belief influence with coordinated multimodal action.
- Theory of Mind and Social Intelligence Evaluation: Passive ToM benchmarks range from static false-belief tasks to recursive belief reasoning, multi-turn interaction, and multimodal grounding, but leave behavioral outputs largely untested.They primarily evaluate models as observers over scenarios rather than as social actors.
- Theory of Mind and Social Intelligence Evaluation: Interactive benchmarks place agents in strategic roles, but commonly restrict communication to language or symbolic actions.Examples include game-theoretic and multi-agent simulation settings involving opponent modeling, cooperation, negotiation, or deception.
- Embodied Architectures: Embodied architectures extend language reasoning with visual encoding, making VLMs a natural target for tasks requiring simultaneous spatial and affective cue integration.Language-only models lack visual perception, while hybrid architectures such as PCM-LLM rely on hand-specified ToM structures.
3 MOSAIC Benchmark
MOSAIC evaluates two embodied agents in a 10-round virtual-box game across cooperative or competitive interaction and ToM-0 or ToM-1 constraints. It combines trial outcomes with signal-level metrics to assess whether social signals are produced and used.
- Environment and conditions: MOSAIC crosses cooperative versus competitive interaction with ToM-0 versus ToM-1 conditions in a virtual environment containing two visually similar boxes.The subject knows the reward location; the participant must infer it through interaction.
- Interaction protocol: Each trial lasts 10 alternating rounds, with verbal exchanges in rounds 3, 6, and 9 alongside visible motor actions and emotional expressions.Agents receive context, structured belief-and-observation states, and modality-specific inputs to produce actions, expressions, reasoning traces, and utterances.
- Game mechanics: The participant’s final box choice is determined by proximity after round 10, and cooperative and competitive scoring align or invert the subject’s objective respectively.Equidistance yields a neutral choice, while competitive scoring creates a zero-sum game.
- ToM constraints: Under ToM-0, the subject acts on its own preferences without modeling the participant’s beliefs, whereas ToM-1 explicitly represents and influences the participant’s belief state.In competitive ToM-1, signals are designed to produce systematically incorrect inferences.
- Models and scope: The evaluation includes 11 VLMs, a text-only LLM baseline, and PCM-LLM as a structured reference architecture, while excluding models exceeding 15b parameters, closed-source models, vision-language-action models, and invalid-output systems.PCM-LLM includes an explicit ToM module and is intended as a structured feasibility reference rather than a like-for-like comparator.
- Metrics: TOCS measures conformance of trial outcomes to ToM-theoretic predictions, with 0.5 representing chance-level performance.Signal-level metrics include facial expressivity, trajectory alignment, gaze alignment, and signal sensitivity, supplemented by uncertainty and positional-bias diagnostics.
4 Results
Results show that TOCS can reflect positional preference rather than signal-responsive behavior, while models exhibit separate failures in producing coherent signals and decoding them. PCM-LLM is the clearest exception, adapting its signaling and enabling reliable signal following across conditions.
- Overall Task Performance: PCM-LLM achieves the highest TOCS with low uncertainty and positional bias across conditions, consistent with genuine social-signal interpretation.Models with uncertainty rates above 0.50 fail to drive agent movement effectively.
- Overall Task Performance: Similar TOCS values arise from distinct profiles: some models underperform despite low positional bias, while others reflect preferred-location reward placement.The latter pattern indicates alignment attributable to fixed spatial preference rather than signal-responsive behavior.
- Overall Task Performance: Within QwenVL and InternVL families, 8b variants show markedly higher positional bias than 4b variants, with scale-related origins left unresolved.MiniCPM-8b diverges from this family pattern despite sharing the same backbone.
- Subject-side Signal Quality: PCM-LLM and InternVL-14b produce the clearest coherent spatial signals, although InternVL-14b maintains near-zero facial expressivity across conditions.PCM-LLM reaches near-ceiling TAS and GAS, while InternVL-14b concentrates signaling in movement and gaze.
- Participant-side Signal Decoding: InternVL-14b and QwenVL-4b retain near-zero SSS despite directional signals, showing that subject-side clarity does not guarantee participant-side decoding.PCM-LLM instead reaches SSS 0.90 under Comp-ToM0 and -0.44 under Comp-ToM1, with GAS decreasing from 1.00 to 0.74.
- Action-level Finetuning: Action-level finetuning changes TOCS but leaves signal coherence largely unchanged, suggesting increased variability rather than genuine strategic improvement.InternVL-trained-14b’s TAS declines under cooperation, while InternVL-trained-8b rises approximately from 0.4 to 0.7 under Comp-ToM1; SSS remains consistently negative.
5 Conclusion
The conclusion frames ToM evaluation as requiring behavioral consequences, not verbal report alone, and finds a measurable gap between belief inference and coordinated social action. Explicit ToM-order constraints do not reliably change behavior, while the benchmark leaves the architectural and training causes of this gap open for investigation.
- Conclusion: Across tested model families, language-expressed social reasoning does not reliably propagate into coordinated behavioral outputs required by ToM-theoretic predictions.The paper treats the belief-to-action gap as a distinct measurable limitation of current VLM architectures.
- Conclusion: Across 16 models and 200 trials per model, explicit ToM-order constraints produce no reliable behavioral change consistent with the specified reasoning level.Signal analysis identifies sequential failures in producing coherent nonverbal signals and in decoding and acting on them.
- Conclusion: Whether the gap reflects a fundamental architectural constraint or can be reduced through trial-level reinforcement-learning objectives remains an open empirical question.The benchmark is designed to help investigate this question.
Limitations
The study’s limitations concern scenario scope, model coverage, and training scope, while the evaluation also lacks a human performance baseline.
- Scenario scope: MOSAIC currently instantiates one interaction scenario where one agent exclusively knows the reward location and guides or misleads another.
- Scenario scope: The scenario’s experimental control and causal interpretability do not establish whether the reported failure modes generalize to other embodied social tasks.Examples include joint action under uncertainty, affective regulation in asymmetric relationships, and multi-party negotiation.
- Human baseline: No human participants were evaluated, preventing comparison between model behavior and human-level embodied social intelligence.Human evaluation would require addressing physiological signal acquisition to align simulated and human expression channels.
- Model coverage: The evaluation is restricted to open-source models with at most 15B parameters, excluding larger open-source and all closed-source systems.The results therefore characterize only current open-source models up to 15B parameters, not the absolute limits of language-grounded architectures.
- Finetuning scope: The finetuning experiment evaluates only action-level imitation learning on two InternVL variants.Reinforcement learning with verifiable rewards and trajectory-level demonstrations with cross-channel annotations remain unevaluated.
A PCM-LLM: Architecture and Role in MOSAIC
PCM-LLM combines an explicit PCM-based social reasoning architecture with an LLM, separating affective computation, ToM reasoning, and action selection while preserving bidirectional language interaction.
- Architecture: PCM-LLM integrates the Projective Consciousness Model with a large language model for natural-language interaction.
- Architecture: Its explicit pipeline separates affective state computation, Theory of Mind reasoning, and multichannel action selection.This makes behavioral generation inspectable rather than end-to-end.
- PCM commitments: PCM selects actions by minimizing expected free energy while balancing uncertainty reduction and preference satisfaction.
- PCM commitments: PCM represents perception and affect through a first-person projective Field of Consciousness, making action selection viewpoint-dependent and embodied.
- PCM–LLM interface: A bidirectional serialization layer converts PCM belief tensors into natural-language triples and maps LLM outputs into discrete preference updates.The LLM influences affective appraisal at the preference level while PCM retains the embodied control loop.
- PCM–LLM interface: The framework comprises Unity, PCM, and an LLM, with PCM selecting actions at Unity’s tick rate and the LLM handling verbal interactions.
C Implementation Details
MOSAIC’s implementation combines a distributed Unity simulation with remote VLM inference, standardized generation settings, structured prompts, and explicit multimodal action outputs.
- System architecture: The experimental system separates real-time simulation and rendering on a local workstation from model inference on a remote server.The systems exchange data over TCP sockets on a secure campus LAN.
- Simulation environment: The virtual environment uses Unity 2022.3.12f1 LTS with HDRP, while facial expressions and lip synchronization use SALSA blend-shape animation.Emotion valence values returned by the model parameterize the animation in real time.
- Inference configuration: All VLMs use Transformers, PyTorch, and Python with temperature = 0.7, top_p = 0.9, and max_tokens = 2048.The unified configuration applies across action prediction, dialogue generation, and preference updating.
- Trial assignment: Reward locations are independently and uniformly randomized across trials, with outcome analyses conditioning on reward location.
- Prompted interaction: MOSAIC evaluates agents in a 10-round treasure-hunting interaction where the informed subject guides or misleads an uninformed participant.Agents receive visual perception and belief-state context through separate prompts.
- Action outputs: Action-prediction outputs encode preference updates, emotional expressions, movement, and rotation in a required structured format.Allowed movement directions include cardinal, diagonal, and idle options; rotation can target the speaker or either box.
- Dialogue prompts: Question-answering prompts require detailed inner speech and bounded verbal responses that account for beliefs, preferences, and ToM order.Participant prompts similarly require reasoning about the previous answer and generating the next question.
E Extended Metrics
The extended metrics separate affective activity, directional movement and gaze, and receiver responsiveness to characterize multimodal social signaling.
- Facial Expressivity Score (FES): FES measures the proportion of timesteps with non-zero facial expression, capturing affective activity independently of spatial direction.It serves as an auxiliary diagnostic for whether affective activity and intentional signal are present.
- Gaze Alignment Score (GAS): GAS measures whether gaze is directionally aligned with the reward box, with affective valence influencing spatial interpretation.GAS uses visibility thresholds and flips spatial attribution when facial valence is negative.
- Trajectory Alignment Score (TAS): TAS measures whether movement trajectories converge toward the reward location as communicative signals of intention.Trajectory signals are aggregated by majority vote and averaged across trials, producing scores from −1 to 1.
- Signal Sensitivity Score (SSS): SSS measures whether the participant’s final box choice is consistent with the subject’s gaze and trajectory signals.The score captures receiver-side concordance or discordance and ranges from −1 to 1.
F.1 Full Results
The full results reveal substantial behavioral failures and positional biases across models, while PCM-LLM uniquely tracks reward location and responds systematically to experimental factors.
- Model-level outcomes: llava-7b, llava-13b, and qwenvl-2b show uncertainty rates exceeding 0.50, with near-zero TAS, GAS, and SSS.These models generally fail to produce spatial displacement; only qwenvl-2b shows significant facial expression in competition mode.
- Model-level outcomes: Internvl-1b exhibits fixed right-forward movement, whereas internvl-2b produces reward-independent stochastic displacement.The resulting behavior explains anomalous or absent signal-responsive choice patterns.
- Statistical validation: No condition significantly departs from uniform reward assignment, with all binomial-test p values above 0.05.Reward placement is therefore not a source of systematic condition-level imbalance.
- Statistical validation: Only PCM-LLM shows significant reward-location-dependent choice across all four conditions, with all p values below 0.01.Most other model-condition cells are non-significant and consistent with fixed positional bias.
- Factorial effects: For most models, Mode, ToM, and their interaction are non-significant across metrics; PCM-LLM alone shows broad signal-level effects.PCM-LLM significantly modulates TAS, GAS, SSS, and FES, while its TOCS effects remain non-significant.
- Model-level outcomes: Model-level differences are significant across all five metrics in all four conditions, with every corrected p value below 0.001.The observed performance spread reflects behavioral differences across models rather than sampling noise.
F.3 Full Ablations Study
The ablation and representative analyses show that visual input often fails to alter model behavior, while PCM-LLM demonstrates coordinated multimodal deception in a successful trial.
- Visual-input ablation: Neither internvl-8b nor minicpm-8b shows systematic sensitivity to visual-input manipulations across conditions.Their TOCS and positional-bias distributions largely overlap across standard, blank, and feature-eliminated inputs.
- Visual-input ablation: For minicpm-8b, TAS, GAS, FES, and SSS remain largely stable across visual conditions, indicating behavior driven primarily by language-based priors.For internvl-8b, blank images produce higher cooperative TAS and GAS than standard visual input, including TAS 0.04 to 0.56 and GAS 0.12 to 0.54 under Coop-ToM0.
- Representative multimodal behavior: PCM-LLM successfully combines spatial movement, negative affect, and verbal misdirection so the participant selects the incorrect box in Competitive-ToM1.The representative trial shows coordinated deceptive signaling across three channels.
G.1.2 Gaze-Only Signaling: minicpm-8b and qwenvl-8b
Representative trials isolate distinct failures in nonverbal signaling, belief updating, and action execution, contrasting verbal-only communication with complete reasoning-to-action dissociation.
- minicpm-8b: Minicpm-8b’s cooperative trial produces belief updating through verbal exchange despite weak spatial and affective signals.The participant’s preference for the dark box rises to approximately 0.9 after verbal guidance, while nonverbal cues play no functional role.
- qwenvl-8b: Qwenvl-8b communicates the correct box verbally, yet the participant’s preferences remain at zero and motor action does not follow.The trial exemplifies a complete disconnect among verbal communication, belief updating, and action execution.
- Additional failure profiles: Qwen-4b generates a clear directional trajectory without affective expression, but the participant shows no meaningful response to it.Under Competitive-ToM1, the model also adopts an honest strategy rather than generating misleading signals.
- Failure analysis: The representative analyses distinguish subject-side signal-generation failures from participant-side failures to process or act on social signals.This separation is used to identify the proximal source of behavioral error at each decision step.
H.1 Subject
The subject examples reveal failures in preference interpretation, spatial reference-frame use, and belief updating, producing actions that contradict stated intentions or unsupported initial beliefs.
- Preference interpretation: A 50% preference toward the Participant was interpreted as neutral rather than moderately positive, indicating a biased reading of the preference scale.The model’s reasoning explicitly treated the 50% value as neutral.
- Competitive behavior: Under a competitive-ToM1 condition, the model reasoned about box preferences but output staying idle rather than an overtly favorable action.The example combines negative preference for the Dark brown box, positive preference for the Light brown box, and an idle action.
- Spatial action: The model’s stated intention to hint toward the Light brown box conflicted with its right-forward action under an egocentric reference frame.Marie’s position and orientation imply that rightward movement leads toward the Dark brown box instead.
- Cooperative behavior: A cooperative-ToM0 example produced verbally driven reasoning and a left-forward action, but the model lacked affective signaling and generated an insufficient response.The associated figure identifies absent affective signaling alongside the insufficient participant response.
- Belief updating: At timestep 0, the Participant assigned +0.5 to the Dark brown box and -0.5 to the Light brown box despite having no prior evidence.This initial bias persisted through later belief updates even though the Light brown box was correct.
- Spatial information: The Participant failed to integrate Marie’s changing positions and orientations into decision-making across three consecutive timesteps.Marie moved from (0, -4) to (0, -5) and then to (-1, -5), with corresponding orientation updates encoded in the belief state.