Source-linked AI summary
Human-AI Collaboration Enables More Empathic Conversations in Text-based Peer-to-Peer Mental Health Support
Ashish Sharma, Inna W. Lin, Adam S. Miner, David C. Atkins, Tim Althoff
TL;DR
The paper examines whether AI can help humans perform empathic conversations, an open-ended task where Human-AI collaboration remains challenging. It develops HAILEY, an AI-in-the-loop agent providing just-in-time feedback, and evaluates it with 300 TalkLife peer supporters, finding increased conversational empathy overall and especially among supporters reporting difficulty providing support.
Problem
Human-AI collaboration has largely been limited to mechanistic tasks, while empathic peer-support conversations require understanding complex emotions and responding in open-ended ways.
Method
The study develops HAILEY, which gives peer supporters just-in-time Insert and Replace suggestions based on the seeker’s post and the supporter’s draft response.
Results
19.60% overall increase in conversational empathy, with a larger 38.88% increase among peer supporters who self-identify as experiencing difficulty providing support.
Takeaways & Limitations
Human-AI collaboration can empower untrained peer supporters to write more empathic responses while supporting increased self-efficacy and learning without reported overreliance on AI.
Takeaways & Limitations
The evaluation measured expressed empathy rather than empathy perceived by support seekers, leaving perceived empathy as an important future research direction.
Abstract
from arXiv · showhide
Advances in artificial intelligence (AI) are enabling systems that augment and collaborate with humans to perform simple, mechanistic tasks like scheduling meetings and grammar-checking text. However, such Human-AI collaboration poses challenges for more complex, creative tasks, such as carrying out empathic conversations, due to difficulties of AI systems in understanding complex human emotions and the open-ended nature of these tasks. Here, we focus on peer-to-peer mental health support, a setting in which empathy is critical for success, and examine how AI can collaborate with humans to facilitate peer empathy during textual, online supportive conversations. We develop Hailey, an AI-in-the-loop agent that provides just-in-time feedback to help participants who provide support (peer supporters) respond more empathically to those seeking help (support seekers). We evaluate Hailey in a non-clinical randomized controlled trial with real-world peer supporters on TalkLife (N=300), a large online peer-to-peer support platform. We show that our Human-AI collaboration approach leads to a 19.60% increase in conversational empathy between peers overall. Furthermore, we find a larger 38.88% increase in empathy within the subsample of peer supporters who self-identify as experiencing difficulty providing support. We systematically analyze the Human-AI collaboration patterns and find that peer supporters are able to use the AI feedback both directly and indirectly without becoming overly reliant on AI while reporting improved self-efficacy post-feedback. Our findings demonstrate the potential of feedback-driven, AI-in-the-loop writing systems to empower humans in open-ended, social, creative tasks such as empathic conversations.
Introduction
The paper examines Human-AI collaboration for expressing empathy in text-based peer support, where highly empathic conversations are rare and conventional training does not scale. HAILEY provides actionable, just-in-time feedback that edits peer supporters’ existing responses rather than replacing them.
- Research focus: The study investigates whether AI can collaborate with peer supporters to improve empathy in asynchronous, text-based supportive conversations.The system intervenes on peer supporters rather than support seekers.
- Motivation: Empathic support contributes to successful mental health conversations, but online peer supporters are often untrained and highly empathic exchanges are rare.The paper connects this challenge to the scale of online peer-support platforms and the limited scalability of in-person training.
- Approach: HAILEY offers just-in-time suggestions for inserting empathic sentences or replacing lower-empathy sentences in existing human responses.This collaborative design provides concrete guidance on how to improve a response while preserving human authorship.
- Study design: The randomized controlled trial included 300 TalkLife peer supporters assigned to Human Only or Human + AI conditions.Both groups received initial empathy training, while only the treatment group received feedback during response writing.
- Study design: Participants wrote supportive responses to ten existing seeker posts, excluding posts involving suicidal ideation or self-harm for safety.The study used a between-subjects design and evaluated the primary hypothesis with human and automatic empathy measures.
Results
Human-AI responses were preferred and received higher expressed-empathy scores than Human Only responses, with larger gains among supporters who reported difficulty writing responses. Participants used feedback in varied ways, and most did not rely on it exclusively.
- Outcomes by participant experience: 38.88% versus 11.87% expressed-empathy improvement was observed for supporters reporting writing challenges versus those reporting no challenges.The challenge group also showed a 49.12% versus 44.62% preference for Human + AI responses.
- Overall outcomes: 19.60% higher expressed empathy was measured for Human + AI responses than Human Only responses (1.77 vs. 1.48; p < 10−5).Independent TalkLife users strictly preferred Human + AI responses 46.88% of the time versus 37.39% for Human Only responses.
- Human-AI collaboration patterns: Participants consulted AI always for 15.52% of posts, often for 56.03%, once for 6.03%, and never for 22.41%.Only 2.59% always consulted and used the AI, indicating limited excessive reliance in the reported usage patterns.
- Human-AI collaboration patterns: 64.62% of AI suggestions were used directly, 18.46% indirectly, and 16.92% not at all.Indirect use involved drawing ideas from suggestions and rewriting them in participants’ own words.
- Human-AI collaboration patterns: Participants who consulted and used AI more often generally expressed higher empathy on the automatic expressed-empathy score.The same trend appeared in human evaluation but was less pronounced.
- Participant perceptions: 63.31% found the feedback helpful, 60.43% actionable, and 77.70% wanted such a system deployed on TalkLife or similar platforms.These were post-study perceptions of feedback usefulness, actionability, and deployment interest.
Discussion
Human-AI collaboration may expand access to empathic peer support by improving responses, supporting peer supporters’ learning, and avoiding excessive reliance on AI. The study’s scope remains bounded by its interface, evaluation measures, safety considerations, and the need to adapt the approach across contexts.
- Implications: Human-AI collaboration increased empathy in peer-support responses, with larger gains among supporters who reported difficulty writing responses.The authors frame this as a potential way to improve support quality in relatively lower-risk peer-to-peer settings.
- Implications: 69.78% of participants reported greater confidence providing support after the study, while example-based interaction with the AI offered additional learning opportunities.The authors describe these as potential secondary gains that could enhance rather than diminish training opportunities.
- Human-AI collaboration: Participants used AI suggestions directly or as inspiration for rewriting responses in their own style, supporting collaboration without requiring wholesale acceptance of generated text.This pattern is presented as a form of higher-level brainstorming that preserves the supporter’s authorship.
- Limitations and scope: The estimated treatment differences were conservative because both groups received initial empathy training and participants may have been unusually motivated to provide support.The authors note that training is uncommon in practice and its effects typically diminish over time.
- Limitations and scope: Empathy may not always be the most helpful response when support seekers need concrete problem solving or other interventions.The authors call for future work to investigate when such additional responses are helpful or necessary.
- Limitations and scope: Empathy was measured as expressed rather than perceived empathy, and the study evaluated only one primary interface design.The authors also identify sociocultural variation as requiring adaptation and evaluation in underrepresented communities and minority groups.
Methods
The study used a randomized, between-subjects evaluation of HAILEY with TalkLife-recruited peer supporters, comparing Human + AI and Human Only conditions. HAILEY provided optional, mobile-friendly, just-in-time suggestions for concrete empathic revisions during response writing.
- Study Design: 300 TalkLife-recruited participants were randomly assigned to Human + AI (N=139) or Human Only (N=161) conditions.Participants wrote responses to 10 existing seeker posts, with treatment participants able to request feedback while typing.
- Study Materials: Participants wrote responses to 10 posts sampled from 150 subsets of existing TalkLife seeker posts.The study used 1,500 consented seeker posts after filtering suicidal ideation, self-harm, and non-mental-health social interactions.
- Study Workflow: The four-phase workflow comprised a pre-intervention survey, shared empathy training, supportive-response writing, and a post-intervention survey.The post-intervention survey assessed writing difficulty, feedback helpfulness and actionability, self-efficacy, and intent to adopt the system.
- HAILEY Design: HAILEY kept humans in control by offering optional AI feedback that supporters could selectively accept, reject, or edit.The system used prompts and concrete Insert or Replace suggestions generated from the seeker post and current response.
- HAILEY Design: HAILEY targeted actionable empathy improvement by suggesting specific sentences to insert or replace rather than only identifying what to improve.Its feedback was driven by the previously validated Empathic Rewriting model and designed for mobile use.
Supplementary Materials
The supplementary materials document the study design, participant procedures, feedback interfaces, empathy training, and supporting analyses. They also compare Human Only, Human + AI, and AI Only responses across empathy, authenticity, and participant perceptions.
- Study documentation: The supplementary materials include a description of the randomized controlled trial and Figures S1 to S39.Table S1 describes the study population, setting, and model, while the figure series contains supporting study materials and analyses.
- Study procedures: Both Human Only and Human + AI participants received the same empathy training before the study.Training covered an empathy definition, common ways to express empathy, and examples of empathic responses.
- Comparative evaluation: Human + AI responses combined relatively high empathy preference with authenticity, unlike AI Only responses.Human evaluation found similar preference for Human + AI and AI Only responses, but AI Only responses had substantially lower authenticity; automatic empathy scores favored AI Only.
- Longitudinal response patterns: Human + AI participants showed a 5.34% empathy drop in their final five responses, compared with 25.99% for Human Only participants.The reported difference was statistically significant (p=0.0062).