Source-linked AI summary

Conceptual Metaphors Impact Perceptions of Human-AI Collaboration

Pranav Khadpe, Ranjay Krishna, Li Fei-Fei, Jeffrey Hancock, Michael Bernstein

arXiv:2008.02311v1cs.HCcs.AI

TL;DR

The paper asks how conceptual metaphors shape people’s experiences and evaluations of conversational AI agents. Across Wizard-of-Oz studies manipulating metaphors along warmth and competence, it finds that low-competence metaphors improve post-use evaluations, while high competence and warmth attract potential users. The paper therefore examines how metaphor choices trade off initial appeal against favorable evaluation and cooperation.

  • Problem

    The paper addresses limited understanding of why conversational AI agents with similar underlying technology receive different user responses, focusing on conceptual metaphors as an expectation-shaping mechanism.

  • Method

    The authors manipulate warmth- and competence-signaling metaphors across Wizard-of-Oz conversational-agent studies using a structured travel-planning task and measure expectations, evaluations, adoption, and cooperation.

  • Results

    Low-competence metaphors improve adoption, cooperation, and evaluations after use, whereas high competence and warmth increase potential users’ likelihood of trying the system.

  • Takeaways & Limitations

    Metaphor choice may help attract users through projected competence and warmth, but lower projected competence may support more favorable evaluations and cooperation after interaction.

  • Takeaways & Limitations

    The study is limited to textual metaphors and conversational AI without strong visual cues, so future work should examine embodied agents and visual metaphors.

Abstract

from arXiv · show

With the emergence of conversational artificial intelligence (AI) agents, it is important to understand the mechanisms that influence users' experiences of these agents. We study a common tool in the designer's toolkit: conceptual metaphors. Metaphors can present an agent as akin to a wry teenager, a toddler, or an experienced butler. How might a choice of metaphor influence our experience of the AI agent? Sampling metaphors along the dimensions of warmth and competence---defined by psychological theories as the primary axes of variation for human social perception---we perform a study (N=260) where we manipulate the metaphor, but not the behavior, of a Wizard-of-Oz conversational agent. Following the experience, participants are surveyed about their intention to use the agent, their desire to cooperate with the agent, and the agent's usability. Contrary to the current tendency of designers to use high competence metaphors to describe AI products, we find that metaphors that signal low competence lead to better evaluations of the agent than metaphors that signal high competence. This effect persists despite both high and low competence agents featuring human-level performance and the wizards being blind to condition. A second study confirms that intention to adopt decreases rapidly as competence projected by the metaphor increases. In a third study, we assess effects of metaphor choices on potential users' desire to try out the system and find that users are drawn to systems that project higher competence and warmth. These results suggest that projecting competence may help attract new users, but those users may discard the agent unless it can quickly correct with a lower competence metaphor. We close with a retrospective analysis that finds similar patterns between metaphors and user attitudes towards past conversational agents such as Xiaoice, Replika, Woebot, Mitsuku, and Tay.

1 INTRODUCTION

The paper examines whether conceptual metaphors shape expectations and evaluations of conversational AI agents, independent of the agents’ behavior. Across studies, low-competence metaphors generally produced more favorable post-use evaluations, while high competence helped attract potential users.

  • The paper analyzes metaphor choice as a design issue for explaining why functionally similar conversational agents can receive very different user responses.The retrospective comparison includes agents such as Xiaoice, Tay, and Mitsuku, but the authors state that metaphors cannot explain the whole story.
  • Conceptual metaphors are proposed as a mechanism that shapes expectations and mediates users’ experiences of AI systems.The paper contrasts metaphors that imply high versus low competence or warmth while holding the interaction itself constant.
  • The authors test warmth and competence as the primary dimensions through which metaphors influence evaluations of AI agents.They measure usability, intention to adopt, and desire to cooperate after interaction with the agent.
  • N = 260 participants evaluated a Wizard-of-Oz conversational agent whose metaphor varied while the wizard remained blind to condition.Participants first encountered the metaphor, then completed a travel-planning task with the agent.
  • Low-competence metaphors increased perceived usability, intention to adopt, and desire to cooperate relative to high-competence metaphors despite human-level agent performance.The authors interpret this pattern as consistent with evaluations depending on the gap between expectations and experience.
  • The studies examine metaphor effects sequentially, from warmth and competence to finer-grained competence levels and initial interest in trying the system.The discussion considers the competing objectives of attracting users and ensuring positive experience and cooperation.

2 RELATED WORK

Prior work shows that expectations, mental models, and design cues influence how people evaluate conversational AI, but does not explain all differences between similar systems. This paper positions conceptual metaphors as an underexamined expectation-shaping mechanism and contrasts assimilation with contrast theory.

  • Pre-use expectations can color evaluations of otherwise identical experiences and continue affecting evaluations after extended interaction.This motivates examining how metaphors establish expectations before users interact with AI systems.
  • AI users often form informal folk theories because performance metrics do not provide a simple, accurate account of probabilistic system behavior.These folk theories serve as guiding beliefs about an AI system’s behavior and goals.
  • Prior studies show that interfaces, voices, and other design aspects affect evaluations beyond an agent’s actual capabilities, while mismatched expectations can lead users to retreat to simpler tasks.Interview research found capability expectations for Siri, Google Assistant, and Alexa can exceed actual performance.
  • The research question asks how metaphors impact evaluations of interactions with conversational AI systems.The paper addresses a gap concerning metaphor effects on expectations and evaluations.
  • Conceptual metaphors are common expectation-shaping devices that help users understand and predict AI behavior across conversational and non-conversational systems.Examples include human roles, a “robotic nose,” and a bot that “Will make you LOL.”
  • Assimilation theory predicts that positive expectations improve evaluations, whereas contrast theory predicts that positive metaphors can backfire when experience falls short.The paper formalizes these alternatives as H1 and H2, respectively.

3 METHODS

The studies use goal-oriented travel planning with a Wizard-of-Oz conversational agent and manipulate metaphors along warmth and competence dimensions. Metaphors are pretested and selected to create controlled treatment conditions, including progressively different competence levels.

  • Collaborative AI task: The study uses a structured vacation-planning task requiring participants to select a hotel package, an outgoing flight, and a return flight.Participants compare options across New York, Berlin, and Paris while considering travel dates and amenities.
  • Collaborative AI task: A Wizard-of-Oz paradigm provides strong, controlled conversational performance while avoiding dependence on a particular AI system’s capabilities.The agent uses a hand-constructed database of hotels and flights to preserve a finite task-focused knowledge setting.
  • Sampling metaphors: Participants are assigned to metaphor treatments defined by low or high warmth and low or high competence, plus a no-metaphor control.The design therefore includes five total conditions.
  • Sampling metaphors: Warmth and competence are treated as the two major axes of social perception, with warmth covering sincerity and good-naturedness and competence covering intelligence, responsibility, and skillfulness.The study operationalizes these dimensions using metaphor ratings on a 5-point scale.

4 STUDY 1: METAPHORS DRIVE EVALUATIONS

Study 1 primes participants with one of four warmth-by-competence metaphors or no metaphor before a travel-planning interaction. Participants report expectations before use and evaluate the system after completing the task.

  • Study 1 uses a between-subjects design with four metaphor conditions—low or high warmth crossed with low or high competence—and a no-metaphor control.Each participant is primed with one metaphor or the control condition.
  • Participants are first shown the assigned metaphor and asked about their expected competence and warmth before interacting with the agent.The pre-use expectation measures occur before the goal-oriented task.
  • Participants then converse with the wizard through a chat widget until completing the travel-planning task.The workflow ends with evaluation and a manipulation check after participants finalize their plans.
  • After the interaction, participants evaluate the AI system and answer questions measuring their experience and the metaphor manipulation.The study also includes debriefing that reveals participants had interacted with a human wizard.

4.2 Measures

The study measured participants’ perceptions, conversational behavior, and demographics using surveys, chat-log analyses, and Wizard-of-Oz interaction procedures.

  • User evaluation measures: Usability was measured with four items covering frustration, ease of use, correction time, and whether the system met requirements.
  • Conversational behavior measures: Participants’ conversational behavior was assessed through language categories, message and conversation length, word counts, and interaction duration.
  • Participants and procedure: Participants interacted with a Wizard-of-Oz AI system recruited through Mechanical Turk and received $4 for an average 15 ± 5-minute survey.
  • Participants and procedure: The study planned for 125 participants based on an 80% power analysis but analyzed 140 after exclusions.

4.4 Wizard-of-Oz manipulation check

The manipulation check excluded participants who suspected the Wizard-of-Oz agent was human, while the expectation analyses tested metaphor effects on pre-use evaluations.

  • Wizard-of-Oz manipulation check: Thirteen participants were excluded after both coders identified suspicion that the agent was human, yielding a 10.4% suspicion level.
  • Pre-use expectations: The analysis used two-way ANOVAs with competence and warmth as categorical predictors and pre-use usability and warmth as outcomes.
  • Pre-use expectations: Competence had a large effect on pre-use usability, with low competence rated 2.41 ± 0.89 versus 3.10 ± 0.86 for high competence.
  • Pre-use expectations: Warmth had a large effect on pre-use warmth, with low warmth rated 2.71 ± 1.12 versus 3.53 ± 1.05 for high warmth.
  • Pre-use expectations: Participants generally expected conversational AI to be highly competent and warm, while metaphors changed those expectations and reduced expectation variance.

4.6 Results: Metaphors impact user evaluations and user attitudes

Metaphor competence strongly shaped post-use evaluations and adoption attitudes, while warmth particularly increased cooperation and exploratory interaction.

  • User evaluations: Low competence produced higher post-use usability than high competence: 4.02 ± 0.72 versus 3.58 ± 1.07.
  • Adoption: Low competence increased intention to adopt, rated 4.25 ± 0.96 versus 3.42 ± 1.33 for high competence.
  • Cooperation: Desire to cooperate was higher with low competence than high competence, increasing from 3.69 ± 0.93 to 4.39 ± 0.67.
  • Conversational behavior: Participants in the high warmth condition asked more questions and explored details such as checked luggage and hotel amenities.
  • Cooperation: High warmth also increased cooperation, with ratings of 4.23 ± 0.81 versus 3.85 ± 0.92 for low warmth.

4.7 Results: Expectations change, but behavior doesn’t

Metaphors changed participants’ evaluations without producing significant LIWC differences in conversation language, while high warmth increased interaction volume and duration.

  • Language behavior: LIWC analyses found no significant language-level differences across conditions, although the authors note that LIWC may miss other language shifts.
  • Interaction behavior: High warmth increased words per conversation from 82 ± 37 to 101 ± 45 compared with low warmth.
  • Interaction behavior: Participants in the high warmth condition spent an average of 4 ± 1.5 minutes longer interacting with the AI system.

4.8 Summary

Competence metaphors produce contrast effects: users are more tolerant of gaps from low-competence systems, while adoption and cooperation decline as projected competence increases. Warmth increases cooperation and interaction duration but does not significantly affect adoption intention.

  • Users are more tolerant of knowledge gaps from low-competence systems but less forgiving when high-competence systems make mistakes.
  • Intention to adopt and desire to cooperate decrease as the competence projected by the AI metaphor increases.
  • High-warmth agents are more likely to elicit cooperation and longer interaction, but warmth does not significantly affect intention to adopt.

5 STUDY 2: THE COMPETENCE-ADOPTION CURVE

Study 2 varied competence metaphors while holding warmth broadly high and measured adoption and cooperation across five treatments plus an unprimed control. Adoption decreased monotonically with expected competence, with the strongest benefit for the lowest-competence metaphor.

  • Procedure: 120 participants were recruited across five metaphor conditions, with 20 participants per condition, plus an unprimed control group.The five metaphors were toddler, middle schooler, young student, recent graduate, and trained professional; all were in the high-warmth half of the space.
  • Procedure: The unprimed system was viewed as roughly as competent as a recent graduate.
  • Results: Intention to adopt decreased monotonically as expected competence increased, while the toddler metaphor produced the greatest beneficial violation of expectations.Projecting more competence than the toddler metaphor incurred an immediate adoption cost, and only the lowest-competence metaphor received a substantial benefit.
  • Results: Intention to adopt differed significantly across metaphor groups, with one-way ANOVA yielding F(4, 114) = 3.701,p = .007.
  • Results: The toddler condition produced the highest desire to cooperate, exceeding middle schooler, young student, recent graduate, and trained professional conditions.The toddler mean was 4.50±0.50; the comparisons were statistically significant for all four other metaphors.
  • Interpretation: Users favored lower-competence agents despite all conditions receiving human-level performance, and adoption declined further when the agent under-performed expectations.

6 STUDY 3: THE COST OF LOW-COMPETENCE METAPHORS

Study 3 examined pre-use reactions without interaction by describing AI services with metaphors and measuring willingness to try, adopt, cooperate, and tolerate errors. Potential users preferred systems projecting high competence and high warmth.

  • Procedure: The study presented metaphor-described AI services without user interaction and measured participants’ likelihood of trying them and using them long term.
  • Procedure: Pre-use cooperation was measured through willingness to cooperate with the system and tolerate its errors.These questions were combined into a pre-use desire to cooperate index.
  • Procedure: 80 new participants evaluated metaphors from four warmth-competence quadrants, with 20 participants exposed to a metaphor from each quadrant.
  • Results: High competence significantly increased trial intention, with average trial indices of 2.95 ± 1.18 versus 1.5 ± 0.83 for low competence.Competence had a significant effect: F(1, 100) = 42.96,p < 0.001,η2 = 0.34.
  • Results: High competence and high warmth both significantly increased participants’ positive pre-use cooperation toward the AI systems.The average pre-use cooperation index was 2.80 ± 1.11 for high competence versus 1.58 ± 0.84 for low competence.
  • Interpretation: Unlike the prior studies, high competence benefited pre-use attraction, while high warmth also increased willingness to try and behave positively toward the service.

7 DISCUSSION

The discussion frames metaphor choice as an expectation-setting design tool with effects that extend beyond conversational agents. It highlights a tension: low competence improves evaluations, while high competence and warmth can attract initial users.

  • Design implications: Metaphors may shape expectations across algorithmic systems, although effects can vary with task, interaction, and context.The paper discusses applications beyond conversational AI, including news feeds, content curation, and recommender systems.
  • Design implications: Studies 1 and 2 found that low competence metaphors produced the highest agent evaluations, whereas Study 3 found they were least likely to be tried.This creates a design trade-off between favorable post-use evaluations and initial adoption.
  • Design implications: Higher warmth was consistently beneficial for attracting cooperation, while competence required balancing user attraction against possible abandonment.The authors suggest lowering competence expectations after interaction begins when using a higher-competence metaphor.
  • Limitations and future work: The study was limited to textual metaphors in a conversational AI without strong visual cues, leaving embodied and visually communicative systems for future work.The authors specifically call for examining visual metaphors and visual signals of competence and warmth.
  • Limitations and future work: Structured travel-planning tasks, 15–20-minute interactions, and possible task novelty limit conclusions about open-ended, prolonged, or familiar use.Future work should examine personal conversations, longer exposure, and effects of prior experience with similar technology.
  • Limitations and future work: The study measured adoption and behavioral intentions, while effects on trustworthiness, task failure, and systems below human-level competence remain open questions.The authors also report only partial support for Assimilation Theory along the warmth axis.
  • Retrospective analysis: Retrospective chatbot comparisons showed that popular systems generally signaled high competence, potentially creating expectations that later interactions could disappoint.Descriptions of Xiaoice, Replika, Woebot, Mitsuku, and Tay were assessed using the study’s warmth-competence measures.
  • Retrospective analysis: The retrospective patterns cannot establish that metaphors alone caused users’ reception because several other variables also affected these systems.The authors present consistency with real-world outcomes as notable but explicitly caution against causal attribution to metaphors alone.

8 CONCLUSION

The paper experimentally examines conceptual metaphors as a causal factor in users’ expectations and evaluations of AI agents. It finds that warmth supports cooperation, while low projected competence improves adoption and cooperation despite designers’ usual preference for high competence.

  • Conclusion: Conceptual metaphors experimentally changed users’ pre-use expectations, post-use evaluations, adoption intentions, and desire to cooperate with AI agents.The conclusion identifies metaphor as a causal factor in these user responses.
  • Conclusion: Warmth increased cooperation, whereas low competence increased adoption and cooperation relative to high competence.This result runs counter to designers’ usual tendency to project high competence to attract users.
Loading 2008.02311v1…