Source-linked AI summary
LLM-Based Selection of Incongruent Verbal and Nonverbal Behavior for Virtual Humans
Parisa Ghanad Torshizi, Stacy Marsella
TL;DR
Existing systems often generate nonverbal behavior from utterance content, despite human behavior also reflecting context and internal states. The paper develops a taxonomy and compares LLM prompting approaches for selecting mismatched behavior, finding that richer context yields richer, context-sensitive outputs whose embodied effects are reliably perceived by human observers.
Problem
Utterance-driven nonverbal generation captures behaviors that illustrate or repeat speech but is limited in modeling contextually shaped mismatches between verbal and nonverbal behavior.
Method
The paper develops a taxonomy of verbal–nonverbal mismatches, compares prompting approaches with progressively richer dialogue context, and evaluates embodied outputs in human-subject studies.
Results
Richer context produced richer, context-driven nonverbal behaviors, while embodied behaviors reliably shifted observers’ perceptions of the virtual human’s attitude and emotion.
Takeaways & Limitations
Context-driven LLM selection supports richer nonverbal behavior in virtual humans, including behavior that reveals information implicit in dialogue and interaction context.
Takeaways & Limitations
The study omitted gestures, postural shifts, and voice prosody because of character-animation implementation issues.
Abstract
from arXiv · showhide
Nonverbal behavior generation systems for virtual agents often take an utterance as input and generate nonverbal behaviors that emphasize or illustrate the content of the verbal channel. However, human nonverbal behavior is shaped by more than the content of the speech. It is also influenced by speaker roles, interpersonal relationships, social context, and the cognitive and emotional states of the interactants. As a result, the nonverbal channel may reinforce, weaken, qualify, or even contradict the verbal channel. It may also reveal internal states that are hidden or only indirectly implied in speech, including emotional "leakage" that may be incidental to the immediate interaction. Modeling this richer relationship between verbal and nonverbal behavior is important for designing virtual agents that exhibit realistic, human-like behavior. It is especially critical in training contexts that require nuanced social interpretation, such as counseling simulations involving virtual patients. Drawing on Ekman's framework of verbal nonverbal relationships, we propose a taxonomy of categories in which mismatches between verbal and nonverbal behavior can occur. We then examine alternative approaches for realizing these behaviors using large language models, focusing on whether LLMs can select contextually appropriate mismatched verbal and nonverbal behaviors from a given dialogue and social interaction context. Finally, we evaluate the resulting behaviors in a human-subject study, assessing whether context-driven nonverbal behavior, when embodied in a virtual human, produces the intended effects on observers.
I. INTRODUCTION
Virtual-human nonverbal generation often follows utterance content, but human behavior also reflects context, relationships, and internal states, including mismatches that convey social information. The paper therefore studies LLM selection of contextually appropriate mismatched behaviors and evaluates their effects when embodied in virtual humans.
- Motivation: Current virtual-human systems often generate nonverbal behaviors from utterances, producing behaviors that illustrate or repeat verbal content but limiting richer verbal–nonverbal relations.This content-driven approach can produce coherent behavior while remaining less expressive than human interaction.
- Motivation: Human nonverbal behavior can repeat, augment, illustrate, accent, or contradict speech, with contradiction conveying meaningful social and communicative information.Nonverbal behavior may reveal hidden negative emotion behind positive speech or maintain warmth during criticism.
- Motivation: Nonverbal cues are less consciously controlled than dialogue and may reveal underlying emotions and attitudes, making them informative in clinical settings.Such cues can indicate a patient’s internal state, psychological processes, and emotional regulation.
- Motivation: Authentic mismatched behavior is especially relevant to virtual-patient counseling training, where trainees must infer unregulated emotional states from nonverbal cues.Related demands arise in negotiation training, social-skills development, and deception-detection scenarios.
- Research aim: The paper explores LLM-based selection of nonverbal behaviors conditioned on interaction context, communicative intent, and mental states.It examines whether LLMs can select mismatched behaviors relevant to a provided dialogue and social interaction context, then evaluates their effects in a human-subject study.
- Research aim: The proposed interface between cognitive, emotional, and dialogue layers and behavior generation must carry more than the utterance when situational and relational factors shape discordant behavior.The paper frames this as a design question for virtual humans.
A. Verbal Nonverbal Relationships
The paper organizes verbal–nonverbal mismatches by their social functions, extending Ekman’s relationships beyond the repeating and illustrating behaviors emphasized by many systems. The taxonomy supplies scenario categories for studying LLM selection of mismatched behavior.
- Framework: Ekman’s framework includes repeating, augmenting, illustrating, accenting, and contradicting relationships between verbal and nonverbal channels.The framework also distinguishes temporal relations such as anticipation, coincidence, substitution, and following.
- Framework: Contradiction is comparatively underexplored even though mismatched channels can carry social and communicative information beyond either channel alone.The related-work context describes a progression from rule-based to data-driven nonverbal generation approaches.
- Taxonomy: The taxonomy groups mismatch contexts by social function and uses them to inspire prompts and construct scenarios for evaluating LLM behavior selection.The taxonomy is not inserted explicitly into prompts to force a mismatch.
- Taxonomy: Irony covers positive verbal content with negative nonverbal behavior and negative verbal content with positive nonverbal behavior.These are described as sarcastic irony and kind irony, respectively.
- Taxonomy: Face maintenance includes strategic politeness and softened criticism, using mismatched behavior to protect the speaker’s or another person’s face.The paper associates these cases with forced positivity or affiliative nonverbals accompanying negative verbal content.
- Taxonomy: Deception and emotion regulation include positive or negative leakage, suppression leakage, and masking of negative words with non-negative nonverbals.Leakage concerns concealed information appearing through nonverbal channels, while suppression leakage concerns negative internal states behind a positive verbal front.
IV. APPROACH
The approach compares prompting strategies that provide progressively richer information, allowing an LLM to decide from context whether verbal and nonverbal behavior should match or mismatch. Eight socially grounded scenarios instantiate the taxonomy for this comparison.
- Prompting approaches: Three prompting approaches vary the information supplied to the LLM, from the utterance alone to dialogue history and dialogue plus context.The approaches extend systems that use only the utterance as input.
- Prompting approaches: All prompts allow context to determine whether verbal and nonverbal behavior matches or mismatches rather than enforcing a mismatch externally.The design tests whether the LLM can automate the mismatch decision.
- Scenarios: Eight short dyadic scenarios represent taxonomy sub-cases involving relationships suggestive of possible verbal–nonverbal mismatch.Examples include workplace sarcasm, playful friendship sarcasm, strategic politeness, softened criticism, and positive or negative leakage.
- Prompting approaches: The prompts ask the LLM to analyze each scenario, describe nonverbal behavior, and report its relation to the verbal behavior.Requested behavior descriptions include action units, gaze, and head movements.
- Scenarios: The dialogue-plus-context prompt begins with a short interaction description before presenting dialogue that establishes a junior employee responding positively to a senior manager’s workload-increasing policy.The example ends with an apparently agreeable utterance from the junior employee.
V. OUTPUT ANALYSIS
The output analysis finds that richer prompting produced more expressive, context-consistent behaviors and better differentiation between verbal and nonverbal information. The selected behaviors were also suitable for influencing human interpretations when embodied in a virtual human.
- Experimental setup: 24 prompts tested eight scenarios under three prompting approaches using Claude Sonnet 4.6 with temperature 0.7 and a 1024-token limit.The experiments were conducted through the Anthropic API.
- Qualitative assessment: Claude’s example output selected restrained smiles, brow tension, gaze aversion, lip pressing, and other timed behaviors that contradicted the speaker’s verbal agreement.The output characterized the relation as contradicting in a junior–senior workplace interaction.
- Qualitative assessment: The qualitative assessment found behaviors expressive and consistent with the supplied situation, with descriptions specifying style, direction, and timing.These properties were considered important for eventual automatic mapping to character animation.
- Prompt comparison: Increasing context increased behavioral richness and enabled Claude to distinguish verbal from nonverbal information.The analysis examined action units, gaze shifts, and contrasts between verbal and nonverbal attitudes or emotions.
- Prompt comparison: Different contrastive conditions showed different distributions of nonverbal behaviors, with marked action-unit differences consistent with expectations.Figure 1 focuses on action units with marked differences in occurrence.
VI. HUMAN SUBJECT STUDY
The study evaluates LLM-selected repeating and contradictory nonverbal behaviors embodied in a virtual human, using two human-subject studies and contrastive stimuli.
- Two human-subject studies evaluated LLM-generated nonverbal behaviors realized in a virtual human and presented as video stimuli.
- The scenarios held Speaker B’s verbal utterance constant while varying nonverbal behavior between repeating and contradicting conditions.
- Positive verbal content was paired with either positive nonverbals in pos-pos or negative nonverbals in pos-neg.
- Negative verbal content was paired with either negative nonverbals in neg-neg or positive, warm nonverbals in neg-pos.
B. Virtual Human Realization
The virtual human realization mapped LLM-selected behaviors onto animation parameters, with additional facial adjustments needed to accommodate the character’s default configuration.
- B. Virtual Human Realization: LLM-selected behaviors were mapped to animation parameters for Action Units and timing relative to words and phrases.
- B. Virtual Human Realization: Figure 2 contrasts a matching pos-pos condition with a mismatching pos-neg condition in the virtual human framework.
- B. Virtual Human Realization: Small AU1 and AU2 activations adjusted the virtual human’s default facial expression because the LLM does not account for the character’s default skeleton and facial configuration.
- B. Virtual Human Realization: The realized stimuli were recorded as videos for participant evaluation.
3) Participants:
Study 1 used within-subject comparisons of matched and mismatched verbal–nonverbal videos to test perceived speaker attitude, while Study 2 independently rated emotions and channel valence.
- 3) Participants:: 49 participants were recruited for Study 1 through Prolific after excluding one incomplete response; participants were aged 18–54.
- 3) Participants:: 43 of 49 participants (87.8%) chose pos-pos over pos-neg, while 46 of 49 (93.9%) chose neg-pos over neg-neg as more positive.
- 3) Participants:: Study 2 had participants rate each of four videos independently, assessing expressed emotions, perceived verbal valence, perceived nonverbal valence, and open-ended interpretations.
E. Hypotheses
The hypotheses predict that LLM-selected nonverbals shift perceived emotion toward nonverbal valence and that intended and perceived nonverbal valence align.
- E. Hypotheses: H2 predicts that LLM-selected nonverbal behaviors shift perceived expressed emotion toward the nonverbals’ valence.
- E. Hypotheses: For positive verbal content, H2.a predicts negative nonverbals produce more negative perceived emotion than positive nonverbals.
- E. Hypotheses: For negative verbal content, H2.b predicts positive nonverbals produce more positive perceived emotion than negative nonverbals.
- E. Hypotheses: H3 predicts alignment between the LLM’s intended nonverbal valence and participants’ perceived nonverbal valence.
- E. Hypotheses: Study 2 summarized positive and negative emotion composites by averaging seven positive and seven negative emotion items, respectively.
3) Analysis of Variance:
Repeated-measures ANOVAs tested whether condition affected perceived emotions and verbal/nonverbal positivity across four pairing conditions. Condition significantly influenced all four dependent variables, with pairwise differences generally favoring positive nonverbal behavior.
- Analysis procedure: The analyses used one-way repeated-measures ANOVAs with Greenhouse-Geisser correction when needed and Bonferroni-adjusted posthoc comparisons after significant omnibus tests.Generalized eta-squared was reported as the effect-size measure.
- Positive emotion: Condition significantly affected positive emotion, F(2.21, 83.82) = 40.04, p < .001, η2 G = .267.The pos-pos condition exceeded pos-neg, while neg-pos exceeded neg-neg in positive emotion.
- Negative emotion: Condition significantly affected negative emotion, F(2.23, 84.56) = 37.22, p < .001, η2 G = .336.Pos-pos had lower negative emotion than pos-neg, whereas neg-neg had higher negative emotion than neg-pos.
- Verbal positivity: Condition significantly affected verbal positivity, F(2.24, 85.15) = 29.14, p < .001, η2 G = .365.Pos-pos and pos-neg did not differ significantly, but neg-pos scored higher than neg-neg (p = .002).
- Non-verbal positivity: Condition significantly affected non-verbal positivity, F(2.66, 101.26) = 21.92, p < .001, η2 G = .310.Both pos-pos versus pos-neg and neg-pos versus neg-neg showed significantly higher non-verbal positivity in the positive-nonverbal condition (both p < .001).
4) Qualitative Analysis:
Participants’ free responses distinguished the intended emotional and attitudinal meanings across matched and mismatched verbal–nonverbal pairings. Positive nonverbals were interpreted as happiness or compassion, while negative nonverbals signaled sarcasm, anger, or disappointment.
- Participant interpretations: Pos-pos was interpreted as happy and non-sarcastic, with the speaker communicating how happy she was.
- Participant interpretations: Pos-neg was interpreted as sarcasm, anger, and aggressiveness because the speaker’s words contradicted her facial expression.
- Participant interpretations: Neg-pos was interpreted as compassion and concern, alongside slight disappointment about the son’s grades.
- Participant interpretations: Neg-neg was interpreted as anger and disappointment without compassion, conveying upset about the son’s grades.
VII. DISCUSSION AND CONCLUSION
The studies found that richer context enabled Claude to select more nuanced nonverbal behavior, and human observers reliably perceived the intended attitude and emotion changes. The approach remains limited by LLM interpretive control, animation-mapping challenges, omitted behavior dimensions, and exploratory evaluation with few stimuli.
- Richer context led Claude to select action units, gazes, and head movements that went beyond illustrating dialogue and revealed information implicit in context and dialogue history.The comparison covered three prompting conditions and eight scenarios.
- Human subject study 1 found that nonverbal behaviors embodied by a virtual human reliably changed participants’ interpretations of the speaker’s attitude.
- Human subject study 2 found that negative or positive nonverbal behaviors reliably shifted perceptions of the virtual human’s emotional valence in the corresponding direction.Nonverbal emotional valence also appeared to shift perceptions of the verbal channel’s emotional valence despite identical verbal content and prosody across conditions.
- The prompting approach cedes control to the LLM to interpret the provided context, while selected behaviors still require mapping to the character animation framework and 3D model.The mapping challenge is especially important for subtle behaviors, including their timing, magnitude, direction, and speed.
- The study omitted gestures, postural shifts, and voice prosody, and its findings should be treated as exploratory because the stimulus set was small.