Source-linked AI summary
From Search Agents to Dissemination Interfaces: Understanding Human Trust in Health Information from Conversational Search
Xin Sun, Rongjun Ma, Xiaochang Zhao, Janne Lindqvist, Jan de Wit, Zhuying Li, Abdallah El Ali, Jos A. Bosch
TL;DR
Trust in health information may depend on both the search agent and the interface delivering it, but this relationship remains underexplored. Across two mixed-methods studies, the paper compares ChatGPT with Google and three LLM dissemination interfaces, finding trust differences across both dimensions.
Problem
Trust in health information across search agents and dissemination interfaces remains underexplored despite the importance of trustworthy health information access.
Method
The paper combines a laboratory study with semi-structured interviews to examine trust across search agents and LLM dissemination interfaces.
Results
Trust varied significantly across search agents and interfaces, with participants trusting ChatGPT more than Google and interface modality influencing LLM-sourced information trust.
Takeaways & Limitations
Trust in conversational health search is shaped by both information source agents and dissemination interfaces.
Takeaways & Limitations
Interaction depth and related factors may have acted as hidden covariates influencing trust formation.
Abstract
from arXiv · showhide
Large Language Models (LLMs) deployed through Conversational User Interfaces (CUIs) are transforming health information-seeking by offering immediate, interactive experiences compared to traditional search engines like Google. However, how trust is influenced by both the types of search agents and the interface used to disseminate the information remains underexplored. This research integrates two mixed-methods studies (lab sessions and interviews) to comprehensively explore trust perceptions in health information across different search agents and dissemination interfaces. In Study 1 (N=21), we investigated trust in health information sourced from ChatGPT and Google across three types of health-related search tasks. Results showed significantly higher trust in health information from ChatGPT, highlighting the promise of LLM-powered conversational search. Building on this, Study 2 (N=20) extended the investigation to explore how the dissemination interface influences trust in LLM-sourced health information by comparing three interfaces: text-based, speech-based, and embodied, all sourcing from the same LLM. Findings revealed significant trust variations across the dissemination interfaces. Interviews from both studies revealed key factors influencing trust in LLM-powered conversational search, including source credibility, participants' search autonomy, and prior knowledge as well as the interaction style and modality. Our findings highlight the potential of LLM-powered conversational search to transform health information-seeking, underscoring the interplay between the credible search agents and the thoughtfully designed dissemination interfaces in shaping trust. These insights are crucial for developing effective, trustworthy LLM-powered health tools to enhance the health information-seeking experience.
1 Introduction
This work examines how search agents and dissemination interfaces shape trust in online health information, addressing underexplored trust mechanisms in LLM-powered conversational search. Using two complementary studies, it compares Google with ChatGPT and text-, speech-, and embodied interfaces to derive trust-related design implications.
- Motivation: Online health information affects health decisions and outcomes, making accuracy, presentation, accessibility, and trustworthiness especially consequential.Trust influences how users evaluate online health advice and whether they act on it.
- Research gap: Trust factors in LLM-powered conversational search remain underexplored, particularly across search agents, dissemination interfaces, and the information they deliver.The work asks whether trust differs between Google and ChatGPT, among text-, speech-, and embodied interfaces, and which factors contribute to trust.
- Approach: The research used two complementary mixed-method studies: Study 1 tested Google and ChatGPT in a within-subject lab study (N=21), while Study 2 isolated dissemination-interface effects.Both studies included follow-up interviews examining motivations and factors behind participants’ trust perceptions.
- Findings: Participants trusted ChatGPT more than Google, but favored simple text-based CUIs over speech- or embodied interfaces because of familiarity and ease of use.Conversational fluency and personalization enhanced trust relative to Google’s list-based presentation, whereas multimodal anthropomorphism did not extend that preference.
- Implications: Human-like modality produced a modality-context mismatch: textual human-like responses fostered trust, while speech and embodied interfaces increased engagement but raised privacy and authenticity concerns.The findings indicate that high-sensitivity health tools should balance conversational engagement with perceived safety, usability, reliability, and transparency.
2 Related work
Prior work distinguishes credibility of health information from trust in the agent delivering it, while showing that search interfaces and interaction cues shape trust. However, trust in LLM-powered conversational health search remains insufficiently understood, particularly beyond performance evaluation.
- Trust and credibility: Credibility concerns information believability, whereas trust concerns reliance on the agent and acceptance of vulnerability and risk.Credibility operates at the information level, while trust in the agent reflects judgments about the provider or system.
- Trust and credibility: Online health trust is shaped by source credibility, content reliability and relevance, and the design and usability of the delivery system.These factors align with technology-acceptance perspectives in which perceived usefulness and usability influence attitudes and acceptance.
- Theoretical frameworks: MATCH frames trust through model attributes, afforded cues, and trust heuristics, while MAIN explains how modality and interface cues alter trust perceptions.MATCH includes competence and reliability, interface design and transparency, and shortcuts such as brand reputation or response fluency; MAIN addresses modality, agency, interactivity, and navigability.
- LLM-powered health search: LLM conversational search shifts health information seeking from retrieval to generation, often obscuring original sources and reducing opportunities for external verification.Users therefore rely more heavily on the agent’s internal competence and benevolence than in web search.
- Research gap: Trust in LLM-powered conversational search remains underexplored, while existing health research has focused primarily on evaluating LLM performance.Conversational fluency may encourage overtrust by masking hallucinations, biases, and limited sourcing transparency.
3 Study 1: Comparison of search agents for health information seeking
Study 1 used a mixed-methods approach combining a within-subjects laboratory study with follow-up semi-structured interviews to compare Google and ChatGPT (GPT-4o) as health-information search agents.
- Study design: The study combined a laboratory-based study with subsequent semi-structured interviews.
- Study design: The laboratory study used a within-subjects design to examine participants’ interactions with search-agent outputs.
- Compared search agents: The comparison involved Google and the LLM-powered agent ChatGPT, specifically GPT-4o.
3.1 Study Methods
Study 1 used a within-subjects lab design with 21 participants comparing Google and ChatGPT across three health-question types. Quantitative trust measures and qualitative interviews assessed participants’ experiences, trust judgments, and search strategies.
- Participants: 21 participants were recruited after a power analysis targeting a medium effect size of 0.25 with α=.05 and 80% power.Participation was voluntary, required informed consent, and participants needed English proficiency and online-search experience.
- Search tasks: Each participant completed six health-information search tasks spanning general, symptom-and-cause, and treatment-related questions.Tasks were selected from an open-source dataset of personal health questions labeled by type.
- Procedure: The within-subjects procedure assigned three tasks to Google and three to ChatGPT, with agent and task order counterbalanced.Participants could reformulate queries and interact freely, while Google tasks required clicking at least one result.
- Measures: Trust in retrieved information was rated after each task, while overall trust in each search agent was assessed after completing its task block.The study also measured general propensity to trust technology and prospective intention to use each agent for future health-information seeking.
- Analysis: Quantitative analyses used assumption tests, repeated-measures ANOVA, paired-samples t-tests, and correlations, alongside thematic analysis of semi-structured interviews.The lab session ended with a 25-minute interview about interaction strategies, comparative trust reasoning, and validation practices.
3.2 Quantitative Findings: Lab Sessions
In lab sessions, participants trusted ChatGPT more than Google for health information and as a search agent. Agent-related trust differences remained consistent across health-question types, while trust correlations varied by agent.
- Participants’ general propensity to trust technology was moderately high (M=3.62, SD=.76 on a 5-point scale).
- Mean trust in health information was higher for ChatGPT (4.05, SD=0.47) than Google (3.77, SD=0.64).
- F(1, 20) = 6.73, p= .017, η2 = .057 indicated a significant search-agent effect, with ChatGPT trusted more than Google.
- Trust did not vary significantly by search task, F(2, 40) = 0.63, p= .480, η2 = .006, nor through agent-task interaction, F(2, 40) = 0.20, p= .777, η2 = .002.
- The search-agent trust difference was 0.27, t(21)=-2.53, p=.02, Cohen’s d=0.55, favoring ChatGPT over Google.
- Trust in information correlated with trust in Google (r(21)=0.63, p=.003) but not ChatGPT (r(21) = 0.09, p= .21).Technology trust correlated with ChatGPT information trust (r(21) = 0.49, p= .02) and ChatGPT use intention (r(21) = 0.48, p= .026), but not Google measures.
3.3 Qualitative Findings: Semi-Structured Interviews
Interviews show that people use online health information cautiously and selectively, often combining sources to guide next steps rather than seeking definitive solutions. Trust is shaped by prior knowledge and experience, autonomy, presentation, interaction style, and personalization.
- How people search for health-related information: Participants searched online mainly to refresh knowledge, learn general information, or guide next steps, selectively adopting advice according to urgency, practicality, and risk.They withheld trust because online information may be inaccurate and did not expect perfect solutions.
- How people search for health-related information: Participants combined forums, social media, official websites, Google Scholar, and search agents, often double-checking multiple sources to validate health information.Some proposed using ChatGPT first for orientation, then continuing the search on Google.
- What factors influence trust in information: Trust depended on prior knowledge and experience: positive outcomes increased trust, whereas misinformation, limited ChatGPT familiarity, commercial content, and filter bubbles weakened it.Participants found ChatGPT useful for quick answers but remained cautious when verification exposed incorrect DOI information.
- What factors influence trust in information: Professional yet understandable language, balanced confidence, logical structure, and appropriate visual cues influenced perceived credibility and trust.Specialized terminology could enhance credibility while reducing understandability, and both uncertain and overly certain answers lowered trust.
- What factors influence trust in information: Search autonomy and human-like interaction shaped trust, while participants also wanted ChatGPT to personalize answers to regional systems, symptoms, allergies, and appointment needs.Google’s control supported independent judgment, whereas ChatGPT’s relevant, context-aware dialogue supported follow-up and personalized responses.
4 Study 2: Comparison of dissemination interfaces for health information seeking
Study 2 examines whether trust in LLM-powered conversational health search changes with the dissemination interface, beyond the text-only setting studied previously. Using the same LLM backend across text-based, speech-based, and embodied interfaces, it isolates how CUI modality modulates trust.
- Study 2 rationale: Study 1’s higher trust in ChatGPT than Google was associated with the conversational and “human-like” interaction, but was limited to text.This left unclear whether trust arose from the LLM’s content generation or could be amplified or hindered by voice or embodiment.
- Study 2 design: Study 2 focuses exclusively on LLM-powered conversational search rather than comparing it with Google.The study builds on insights from the MAIN model concerning information modality heuristics.
- Study 2 design: Holding the information source constant with the same LLM backend isolates the effects of text-based, speech-based, and embodied dissemination interfaces.The comparison tests how CUI modalities modulate trust in generative health information.
4.1 Study Methods
Study 2 recruited 20 participants to compare trust across three LLM-powered dissemination interfaces using standardized health-search tasks and mixed quantitative–qualitative methods. The study used the same GPT-4o model across text, speech, and embodied conditions, with trust measured after each search and interviews conducted afterward.
- Interfaces: Three interfaces—a text-based chat, speech-based dialogue, and embodied system—used the same GPT-4o model to ensure generation consistency.The embodied interface used a custom physical body with simple facial expressions to isolate physical presence without hyper-realistic anthropomorphism.
- Tasks and procedure: 75 questions covered general, symptom, and treatment-related health searches, aligning with Study 1 for methodological consistency.Participants completed three tasks per interface in counterbalanced order, totaling nine tasks, and could ask follow-up questions until satisfied.
- Measures: Trust was measured after each search, while usability and intention to use were assessed after each interface condition.Participants also completed eHealth and AI literacy questionnaires before the lab study and joined 15-minute semi-structured interviews afterward.
- Analysis: A Mixed Linear Model analyzed trust differences across interfaces and search-task types while controlling for usability, alongside thematic interview analysis.Statistical assumptions were checked with Shapiro-Wilk and Bartlett’s tests, and Pearson correlations examined variable relationships.
4.2 Quantitative Findings: Lab Sessions
Lab-session results showed high literacy and favorable technology attitudes, with text-based interfaces receiving the highest usability, familiarity, intended use, and trust ratings. Trust in health information correlated with interface trust and usability, while MixedLM revealed significant trust differences involving embodied interfaces.
- Participant characteristics: Participants reported favorable technology attitudes (PPT M=3.87, SD=.33), moderately high eHealth literacy (M=3.68, SD=.47), and high AI literacy (M=3.70, SD=.26).These descriptive results indicate generally positive technology attitudes and relatively high health- and AI-related literacy.
- Interface evaluations: Text-based interfaces had the highest usability (M=4.05, SD=.43), familiarity (M=3.60, SD=.75), and intention to use for health searches (M=3.55, SD=.86).Speech-based and embodied interfaces scored lower on these measures.
- Trust across interfaces: Trust in health information was highest for text-based (M = 4.19, SD=.42), followed by speech-based (M=4.13, SD=.46) and embodied interfaces (M=4.00, SD=.48).Trust in the interfaces themselves followed the same ordering: text-based (M=3.73, SD=.37), speech-based (M=3.57, SD=.46), and embodied (M=3.56, SD=.44).
- Trust across tasks: Trust levels were relatively consistent across search-task types in all three interfaces, indicating similar trust regardless of interface used.The task-level pattern is reported in Figure 7.
- Correlational findings: Across interfaces, trust in health information was significantly related to trust in the interface, with usability also significantly associated for text-based, speech-based, and embodied interfaces.Correlations between information trust and interface trust were r(20)=.58, p<.01 for text-based, r(20)=.52, p<.05 for speech-based, and r(20)=.76, p<.01 for embodied interfaces.
- Mixed-model findings: MixedLM found significant information-trust differences between text-based and embodied interfaces (𝛽= .189, 𝑝= .035) and speech-based and embodied interfaces (𝛽= .180, 𝑝= .002).The reported differences were largely influenced by usability, which significantly predicted trust.
4.3 Qualitative Findings: Semi-Structured Interviews
Interviews showed that trust in health information depends on interface familiarity, usability, presentation and modality, with text-based formats often easier to process. Participants recommended source verification, context-aware interaction, multimodal presentation and personalized responses to strengthen trust.
- Trust factors: Prior experience and interface familiarity shaped trust, with frequent Google or professional-health-site users finding text-based interfaces more familiar and trustworthy.Similarity to social media, health professionals and professional literature further increased trust in text-based interfaces.
- Trust factors: Usability strongly affected trust: participants favored text-based UIs for their simplicity, while usability was especially important for less familiar speech and embodied interfaces.The interviews linked usability to trust mediation in text-based interfaces and to trust-building in speech and embodied formats.
- Trust factors: Presentation style influenced trust because numbered formats suited text but could overload speech-based delivery, whereas storytelling or opening summaries could improve engagement and trust.Participants associated vocal information overload with cognitive overload and forgetfulness.
- Trust factors: Text was easier to process than vocal information because reading required less effort, supported cross-comparison and revisiting, and enabled sharing and clarification.Embodied interfaces additionally required processing verbal information alongside facial and other non-verbal cues, increasing cognitive load.
- Recommendations: Participants recommended linking claims to original sources, confirming query understanding, asking contextually relevant follow-up questions, and personalizing information to symptoms, allergies and regional medical systems.They also recommended combining text, images and videos, and synchronizing embodied speech with expressive features to improve comprehension, usability and trust.
5 Discussion
Trust in health information varied across search agents and dissemination interfaces, with participants generally trusting ChatGPT more than Google and text-based interfaces more than speech-based or embodied ones. Trust was shaped by source credibility, search autonomy, prior knowledge, familiarity, usability, and interaction modality.
- Study 1: Participants reported higher trust in ChatGPT than Google, associating ChatGPT’s conversational, rapid, personalized responses with feeling heard, understood, and receiving relevant expertise.ChatGPT’s intuitive and user-friendly interface may have further supported these perceptions.
- Study 1: Trust in Google as an agent significantly correlated with trust in its health information, whereas ChatGPT trust did not significantly correlate with trust in generated content.This source dissociation suggests users evaluated ChatGPT’s synthesized responses through presentation, prior experience, logical flow, and professional language rather than agent reputation alone.
- Study 2: Trust ratings were significantly higher for text-based interfaces than speech-based or embodied interfaces, despite the latter offering more natural interactions.Participants valued text for precise cross-referencing and familiarity, especially in high-stakes health scenarios, and interface trust directly influenced information trust.
- Trust drivers: Prior knowledge moderated search autonomy: knowledgeable users valued Google’s source verification, whereas users with limited knowledge favored ChatGPT’s synthesized summaries as a starting point.In LLM-based search, autonomy was also expressed through iterative follow-up questions, while prior knowledge reduced the need for extensive autonomous searching.
- Trust drivers: Familiarity strongly favored Google and text-based interfaces because they aligned with established health-information-seeking habits, while speech and embodied interfaces added cognitive load.The discussion recommends aligning emerging technologies with users’ existing search habits to build familiarity and trust.
- Design implications: Usability supported trust through either cognitive ease or process transparency: ChatGPT users valued quick answers, whereas Google users valued transparent source cross-referencing.Multimodal richness could increase engagement but also cognitive load, usability problems, and errors such as tone misinterpretation or synchronization issues.
6 Limitations and future work
The study’s limitations concern its narrowly selected search agents and health-information focus, controlled study design, unassessed response accuracy, speech-recognition implementation, and limited demographic diversity. Future work should test hybrid technologies, broader domains, controlled and long-term behaviors, accuracy-related trust, improved recognition, diverse users, and changing AI systems.
- Search scope and generalizability: The study examined only Google and ChatGPT, while hybrid tools combining web browsing and LLM capabilities warrant investigation across diverse information domains.Examples include Bing and ChatGPT Search; the health-information focus may limit generalizability.
- Study design and interaction: Lab sessions and interviews may not capture spontaneous real-world searching, while varying interaction depths could act as hidden covariates influencing trust.Google behavior ranged from zero-click searches to deep link navigation, and ChatGPT use ranged from single-turn queries to multi-turn dialogues.
- Study design and interaction: Future research should use longer-term and controlled designs to isolate how interaction volume affects trust during naturalistic searching.Suggested controls include fixing the number of turns or clicks.
- Accuracy and reliability: The study measured user perceptions without evaluating factual accuracy, hallucination rates, or their relationship with trust in LLM responses.Future work should examine these relationships, especially as web-searching tools aim to improve reliability.
- Implementation and future validation: Speech-recognition errors appeared to have minimal impact, but future studies could integrate more advanced recognition models and revalidate trust perceptions as AI advances.Participants resolved minor errors through rephrasing, and interviews raised no concerns.
- Demographic diversity: Participants were primarily European students and digitally literate users, limiting generalizability to other demographic groups and users with lower digital literacy.Subsequent research should examine how demographic groups differ in their perceptions and trust of online health information.
7 Conclusion
The research shows that trust in health information is shaped by both the search agent and dissemination interface. ChatGPT received higher trust than Google, while text-based CUIs were preferred over speech-based and embodied interfaces for familiarity and ease of use.
- Search agents and interfaces: The research compared Google and ChatGPT as search agents and text-, speech-, and embodied LLM-powered conversational user interfaces.These comparisons covered both information sources and dissemination modalities.
- Study 1: Participants trusted information from ChatGPT more than Google, influenced by prior experience, presentation style, and interaction mode rather than query type.The reported drivers included prior experience, information presentation style, and interaction mode.
- Study 2: Interface modality further influenced trust, with text-based CUIs preferred for their familiarity and ease of use.This result concerned LLM-sourced information disseminated through different interface modalities.
- Study 2: Speech-based and embodied interfaces enabled natural verbal interaction but raised concerns about privacy and authenticity.The conclusion identifies both an interaction benefit and trust-related concerns for these modalities.
- Implications: Trust was influenced by both source agents and dissemination interfaces, informing future health AI tools designed to be reliable, informative, trusted, and tailored to diverse needs.The findings are presented as a foundation for future health AI tool development.