Source-linked AI summary

Connection-Coordination Rapport (CCR) Scale: A Dual-Factor Scale to Measure Human-Robot Rapport

Ting-Han Lin, Hannah Dinner, Tsz Long Leung, Bilge Mutlu, J. Gregory Trafton, Sarah Sebo

arXiv:2501.11887v1cs.ROcs.HC

TL;DR

Human-robot interaction research lacked a validated scale for measuring rapport across contexts and perspectives. The authors developed and evaluated the 18-item CCR Scale across video-based and in-person studies, finding reliable two-factor structure, acceptable fit, better fit than the Gratch Rapport Scale, and higher rapport with a responsive robot.

  • Problem

    HRI research lacked a validated instrument for measuring rapport across different contexts and from both first-person and third-person perspectives.

  • Method

    The authors developed the CCR Scale through item generation and factor analysis, then evaluated it with new videos, compared it with the Gratch Rapport Scale, and validated it in an in-person interaction study.

  • Results

    The 18-item CCR Scale showed two factors, strong reliability, acceptable confirmatory fit, better fit than the Gratch Rapport Scale, and significantly higher rapport ratings for responsive than unresponsive robots.

  • Takeaways & Limitations

    The CCR Scale can measure human-robot rapport from both first-person and third-person perspectives across varied interaction scenarios.

Abstract

from arXiv · show

Robots, particularly in service and companionship roles, must develop positive relationships with people they interact with regularly to be successful. These positive human-robot relationships can be characterized as establishing "rapport," which indicates mutual understanding and interpersonal connection that form the groundwork for successful long-term human-robot interaction. However, the human-robot interaction research literature lacks scale instruments to assess human-robot rapport in a variety of situations. In this work, we developed the 18-item Connection-Coordination Rapport (CCR) Scale to measure human-robot rapport. We first ran Study 1 (N = 288) where online participants rated videos of human-robot interactions using a set of candidate items. Our Study 1 results showed the discovery of two factors in our scale, which we named "Connection" and "Coordination." We then evaluated this scale by running Study 2 (N = 201) where online participants rated a new set of human-robot interaction videos with our scale and an existing rapport scale from virtual agents research for comparison. We also validated our scale by replicating a prior in-person human-robot interaction study, Study 3 (N = 44), and found that rapport is rated significantly greater when participants interacted with a responsive robot (responsive condition) as opposed to an unresponsive robot (unresponsive condition). Results from these studies demonstrate high reliability and validity for the CCR scale, which can be used to measure rapport in both first-person and third-person perspectives. We encourage the adoption of this scale in future studies to measure rapport in a variety of human-robot interactions.

I. INTRODUCTION

Human-robot rapport reflects mutual understanding and interpersonal connection, but existing HRI measures lack a validated, broadly applicable instrument. The authors introduce the CCR Scale to measure rapport across contexts and perspectives.

  • Existing HRI measures often assess related constructs rather than rapport directly, including attention, positivity, coordination, closeness, satisfaction, and likeability.
  • Prior rapport scales commonly use reverse-coded or context-specific items, reducing response accuracy or limiting applicability across situations.
  • Existing rapport scales also focus on first-person judgments, omitting rapport that observers can assess from a third-person perspective.
  • The CCR Scale is designed as a validated instrument for measuring rapport from both first-person and third-person perspectives in varied human-robot interactions.
  • Rapport is defined as mutual understanding and interpersonal connection developed through interaction.

III. INITIAL SCALE ITEM GENERATION

The authors generated candidate rapport items by combining dictionary and literature sources with public definitions and perceptions of rapport. Delphi grouping and internal review reduced the initial pool to 27 positive interaction items for Study 1.

  • Candidate items were drawn from dictionary definitions, rapport literature, HRI literature, and public descriptions of rapport.
  • Study 0 asked 51 Prolific participants to define rapport and describe interactions characterized by low or high rapport.
  • The Delphi method grouped similar themes and produced 67 candidate scale items.
  • Multiple internal-review rounds reduced the pool to 27 items representing positive interaction characteristics.
  • Study 1 used an online between-subjects design in which participants watched one human-robot video and rated the candidate items.

1) Participants:

Study 1 used four short human-robot interaction videos spanning different contexts and robot characteristics. Participants rated 27 items, and exploratory factor analysis supported a reliable two-factor structure.

  • Materials (Videos): Four videos represented Exercise, Service, Companion, and Argument interactions selected for variation in rapport, context, and robot characteristics.
  • Materials (Videos): The videos were trimmed to under one minute to minimize participant fatigue.
  • Participants: Participants rated 27 items on a five-point Likert scale after being randomly assigned to one video.
  • Results: Exploratory factor analysis used a two-factor promax-rotated solution, retaining items with loadings of at least 0.6 and no cross-loadings above 0.3.
  • Results: The resulting Connection and Coordination factors contained 13 and 6 items, respectively, with high internal reliability (α = 0.96, ωtotal = 0.97).

C. Discussion

Study 2 evaluated the CCR scale on new human-robot interaction videos and compared it with the Gratch Rapport Scale. The videos covered varied robot contexts and behaviors, supporting assessment across scenarios.

  • The CCR scale’s rapport score averages the Connection and Coordination factor scores, which may also be examined separately.
  • Study 2 evaluated the 19-item CCR scale against the Gratch Rapport Scale using a new set of human-robot interaction videos.The online within-subjects study recruited 201 participants after attention-check exclusions.
  • The videos represented healthcare assistance, holding hands, pet companionship, and conversation, thereby varying interaction context and robot behavior.

3) Materials (Scale Items):

Study 2 refined and evaluated the CCR scale by removing a context-specific item and testing its factor structure, reliability, predictive fit, and agreement with participant rankings. Both scales recovered the same video ordering, while the CCR model fit the ranking data better.

  • Materials (Scale Items):: The item “Deep conversation” was removed because the Pet video lacked verbal communication, making the scale more generalizable across robot contexts.
  • B. Results: The CCR scale showed a good CFA fit on most indices, with CFI = 0.997, TLI = 0.996, and SRMR = 0.046, while RMSEA = 0.09 indicated moderate fit.
  • B. Results: The CCR scale was highly reliable (α = 0.97, ωtotal = 0.97) and strongly correlated with the Gratch Rapport Scale (R = 0.84, p < 0.001).The Gratch Rapport Scale was also highly reliable (α = 0.89, ωtotal = 0.94).
  • B. Results: The CCR ordinal-regression model fit the participant ranking data better than the Gratch model, with AIC = 1874.23 versus AIC = 1914.33.Both models were significantly better than chance, and the CCR model was significantly preferred by Nagelkerke Pseudo R2.
  • B. Results: Both scales accurately reproduced the participant ranking from lowest to highest rapport: Healthcare < Hold Hands < Conversation < Pet.

C. Discussion

Study 3 extended validation from video-based third-person ratings to first-person, in-person human-robot interaction. It replicated a responsive-versus-unresponsive robot design using a different robot platform and response modality.

  • C. Discussion: Study 3 addressed whether the CCR scale could measure rapport from the first-person perspective during in-person human-robot interaction.
  • C. Discussion: Participants disclosed a personal concern across three messages while interacting with either a responsive or unresponsive robot.
  • C. Discussion: The responsive robot nodded and delivered personalized positive template-based verbal responses after each participant message, whereas the unresponsive robot provided no gestures or responses.
  • C. Discussion: The replication used a Misty II robot rather than the non-humanoid Travis robot used in the prior study, with Misty combining head nodding and verbal responses.

1) Participants:

Study 3 recruited 44 participants for an in-person comparison of responsive and unresponsive robot conditions, using the CCR scale and related measures.

  • Participants: 44 participants were recruited for Study 3, with 21 participants per condition targeted through an a priori power analysis.The analysis assumed power of 0.8, effect size d = 0.8, and p < 0.05.
  • Measures: The study measured perceived rapport with the 18-item CCR scale and also assessed responsiveness, sociability, competence, attractiveness, and companionship desire.The CCR scale was used alongside measures from Birnbaum et al..
  • Procedure: Participants disclosed a current problem, concern, or stressor to the robot across three messages after receiving study instructions.The experiment used a cover story about testing a speech-comprehension algorithm and asked participants to signal completion of each message.
  • Analysis: Rapport correlated strongly with perceived robot responsiveness (R = 0.75, p < 0.001).Condition differences were analyzed with t-tests, with results summarized in Table III.

1) Connection-Coordination Rapport (CCR) Scale:

Study 3 showed high CCR reliability and higher rapport ratings for responsive than unresponsive robots, with both factors contributing to the difference.

  • Results: The CCR scale showed high internal consistency, with α = 0.95 and ωtotal = 0.96.The scale comprised Connection and Coordination factors whose ratings were also examined separately.
  • Results: Responsive-condition participants rated overall rapport significantly higher than unresponsive-condition participants.The comparison is shown in Figure 5, with error bars representing one standard error from the mean.
  • Results: Responsive-condition participants rated both Connection and Coordination significantly higher than unresponsive-condition participants.Figure 5 compares the two factor scores and the overall CCR score across conditions.
  • Related measures: Participants in the responsive condition also reported significantly higher responsiveness, sociability, competence, attractiveness, and desire for robot companionship.These findings replicated results from Birnbaum et al..
  • Participant’s Self-disclosure: Three coders rated participant behavior from videos, and self-disclosure ratings showed no significant difference between conditions despite high inter-rater reliability (α = 0.91).Self-disclosure ratings ranged from brief, shallow statements to extensive disclosures with deep emotional or personal insights.

C. Discussion

The authors present the psychometrically validated 18-item CCR Scale, whose Connection and Coordination factors measure human-robot rapport across first- and third-person perspectives. Responsive robots elicited significantly greater rapport than unresponsive robots, supporting the scale’s construct validity.

  • C. Discussion: Participants rated significantly greater rapport with the responsive robot than with the unresponsive robot, confirming the CCR Scale’s construct validity.The finding supported the hypothesis that greater robot responsiveness is associated with deeper perceived rapport.
  • C. Discussion: Figure 5 shows similarly high Connection and Coordination ratings for responsive interactions, whereas Connection fell below Coordination for unresponsive interactions.The authors attribute the smaller Coordination decrease partly to immediate robot responses preserving turn-taking, while missing gestures and personalized speech more strongly reduced Connection.
  • C. Discussion: The 18-item CCR Scale contains Connection and Coordination factors identified through exploratory factor analysis and supported by confirmatory factor analysis.The factors align with prior rapport theory by grouping positivity separately from mutual attentiveness and coordination.
  • C. Discussion: The CCR Scale uses brief, forward-coded, broadly applicable items and supports evaluating human-robot relationships from both first-person and third-person perspectives.The authors also suggest using Wizard-of-Oz videos to evaluate robot designs before substantial prototyping.
  • C. Discussion: Future validation should test the CCR Scale with other pair types, group rapport, and longer-term interactions beyond the short-term human-robot pairs studied here.The stated targets include human-human, human-virtual-agent, robot-robot, and multi-person interactions.
  • C. Discussion: The authors conclude that the CCR Scale is psychometrically validated and encourage its adoption in future human-robot interaction research.

I. SUPPLEMENTAL MATERIALS

The supplemental materials document the study instruments, candidate-item decisions, rapport comparison scale, and Wizard-of-Oz response materials. They also provide the survey wording, response scales, and measures used to assess robot responsiveness and related perceptions.

  • I. SUPPLEMENTAL MATERIALS: The supplementary materials include the initial scale-item source table and explain that idiomatic items such as “Being in sync” were rejected as difficult to understand or translate.The listed rejected items include “Having chemistry,” “A bond,” “Being on the same wavelength,” and “Clicking.”
  • I. SUPPLEMENTAL MATERIALS: Study 2 used a modified Rapport Scale 4 from Gratch et al., including connection, mutual understanding, warmth, respect, closeness, and reverse-coded distance or coldness items.The supplementary wording states that bracketed text was modified for the study.
  • I. SUPPLEMENTAL MATERIALS: Study 3 used a Wizard-of-Oz message bank with responsive and unresponsive robot-response conditions.Responsive messages could be selected and customized, whereas unresponsive responses had one fixed option per message and could not be customized.
  • I. SUPPLEMENTAL MATERIALS: The responsive message bank contained empathic and situation-specific responses, while the unresponsive condition instructed participants to continue or contact the experimenter.Examples include “That’s really tough” in the responsive condition and “Please go on to the next part” in the unresponsive condition.
  • I. SUPPLEMENTAL MATERIALS: Supplementary surveys measured perceived robot responsiveness, sociability, competence, attractiveness, and desire for companionship using mostly five-point response scales.The responsiveness items included listening, understanding, shared wavelength, interest, and responsiveness to needs; attractiveness used a seven-point scale.
  • I. SUPPLEMENTAL MATERIALS: The supplemental materials also list perceived-attractiveness and companionship items and explain that survey wording was adapted from Birnbaum et al. with the robot renamed Misty.Companionship items asked whether participants wanted Misty to keep them company during stressful events or when alone.
Loading 2501.11887v1…