Source-linked AI summary
Tell Me About Yourself: Using an AI-Powered Chatbot to Conduct Conversational Surveys with Open-ended Questions
Ziang Xiao, Michelle X. Zhou, Q. Vera Liao, Gloria Mark, Changyan Chi, Wenxi Chen, Huahai Yang
TL;DR
Traditional survey platforms lack interactive support for motivating and guiding quality responses to open-ended questions. The paper compares an AI-powered chatbot survey with a typical Qualtrics survey in a field study of about 600 participants and over 5200 free-text responses. The chatbot elicited significantly more relevant, specific, and clear responses and greater participant engagement, while motivating design implications for conversational surveys.
Problem
Existing survey platforms lack interactive features for automatically interpreting, probing, and guiding responses to open-ended questions.
Method
The study compared a chatbot-driven survey with a Qualtrics survey in a field study involving about 600 participants and analyzed over 5200 free-text responses.
Results
The chatbot elicited significantly more relevant, specific, and clear open-ended responses than the Qualtrics survey and encouraged greater participant engagement.
Takeaways & Limitations
The findings suggest chatbot surveys can support effective collection of free-text responses, with active listening and early intervention as design implications.
Takeaways & Limitations
The study has limitations involving study operations and the scope of its results.
Abstract
from arXiv · showhide
The rise of increasingly more powerful chatbots offers a new way to collect information through conversational surveys, where a chatbot asks open-ended questions, interprets a user's free-text responses, and probes answers whenever needed. To investigate the effectiveness and limitations of such a chatbot in conducting surveys, we conducted a field study involving about 600 participants. In this study with mostly open-ended questions, half of the participants took a typical online survey on Qualtrics and the other half interacted with an AI-powered chatbot to complete a conversational survey. Our detailed analysis of over 5200 free-text responses revealed that the chatbot drove a significantly higher level of participant engagement and elicited significantly better quality responses measured by Gricean Maxims in terms of their informativeness, relevance, specificity, and clarity. Based on our results, we discuss design implications for creating AI-powered chatbots to conduct effective surveys and beyond.
Q. VERA LIAO, IBM Research AI, USA
The paper identifies its authors and frames the work within human-centered computing, intelligent agents, conversational agents, chatbots, surveys, and open-ended questions.
- The paper lists Changyan Chi, Wenxi Chen, and Huahai Yang among its authors.
- Its stated computing areas include human-computer interaction and intelligent agents.
- The work concerns conversational agents, chatbots, surveys, and open-ended questions.
- Figure 1 shows a screenshot of a chatbot survey used in the study.
1 INTRODUCTION
Online surveys are convenient but struggle with fatigue and open-ended responses, while conversational chatbots may add interaction that improves engagement and response quality. The paper addresses this gap through a field comparison of chatbot and traditional surveys.
- Survey context: Online surveys provide broad, flexible, automated data collection, but survey fatigue increases as surveys lengthen.Participants may spend 5 minutes on a 10-question survey but 10 minutes on a 30-question survey.
- Survey context: Open-ended questions offer freely phrased, deeper insights but require extra effort to formulate and type.
- Survey context: Existing survey platforms generally lack automated feedback and probing that could motivate and guide higher-quality open-ended responses.
- Conversational surveys: A conversational survey chatbot asks open-ended questions, probes answers, and handles social dialogues.
- Conversational surveys: Chatbots may personalize questions, support social interaction, provide feedback, probe responses, and manage complex exchanges.
- Research questions: The study asks how chatbot-driven surveys differ from traditional online surveys in response quality and participant engagement.
- Study overview: A field study compared the holistic effects of an AI-powered chatbot and a typical online survey using about 600 participants and over 5000 responses.
- Contributions: The paper reports higher-quality responses and engagement and discusses design implications for effective chatbot surveys.
2 RELATED WORK
This section situates the study at the intersection of conversational AI, information elicitation, and survey-interface research. It identifies unresolved concerns about imperfect conversational capabilities and introduces a comparison focused on free-text response quality and participant engagement.
- Conversational AI and information elicitation: Conversational AI has been studied for roles including personal assistance, tutoring, customer service, job interviewing, and companionship.Related work also uses conversational agents to elicit information for recommendations, decision support, psychotherapy, voting, interviews, and longitudinal studies.
- Conversational AI and information elicitation: Conversational interfaces offer natural expression and flexibility, but technical difficulties remain in interpreting diverse language and managing complex, nonlinear conversations.These shortcomings may affect survey participants and survey results when users provide open-ended responses.
- Conversational surveys: Prior studies generally examined chatbot feasibility for specific elicitation tasks, whereas this work investigates a chatbot as a general surveying tool.Earlier examples focused on domains such as student team-building preferences or job-candidate information.
- Conversational surveys: Unlike typical noninteractive, nonadaptive online surveys, chatbots can ask questions, probe responses, and support interactive information elicitation.The section contrasts fixed survey prompts with conversational follow-up and clarification capabilities.
- Open questions: Existing work had not examined how imperfect chatbot conversation capabilities affect open-ended information elicitation, including response quality and satisfaction.The section also notes concerns about trust, privacy, and sensitive-information sharing in chatbot-mediated elicitation.
- Study contribution: The study compares conversational and traditional GUI surveys by measuring free-text response quality and participant engagement with content-based metrics.It frames this focus as distinct from prior comparisons centered on response rate or broader survey effects.
3 STUDY METHOD
The study uses a between-subjects field experiment to compare an AI-powered chatbot survey with a typical form-based survey on collected-information quality and participant engagement.
- Study design: The field study compares two survey methods: an AI-powered chatbot survey and a typical form-based survey.The comparison measures both the quality of collected information and participant engagement.
3.1 Study Background
The study was conducted with a global market research firm to preserve ecological validity and practical value while examining gamers’ opinions of recently released video game trailers.
- Study context: A global market research firm partnered in the study to support ecological validity and practical value.The firm specializes in discovering customer insights for the entertainment industry, including game companies and movie studios.
- Study context: The field study gathered gamers’ opinions about two video game trailers recently released at E3 2018.The entertainment-industry setting aligned the research study with the collaborator’s market-research goals.
3.2 Study Platform
The study compares Qualtrics’s sequential text-box survey with a Juji chatbot survey that supports customizable, rich conversational interaction. The design examines the holistic effect of these capabilities while acknowledging that chatbot interpretation remains imperfect.
- Study platforms: The study was implemented on two platforms to compare a chatbot survey with a typical form-based survey.The platforms were Qualtrics for the form-based condition and Juji for the conversational condition.
- Study platforms: Qualtrics presents one open-ended question and text box at a time, requiring participants to submit each answer before proceeding.This represents the sequential, nonconversational baseline used in the comparison.
- Study platforms: Juji lets survey creators enter questions and their order through a GUI, then automatically builds a chatbot with capabilities for open-ended dialogue and user digressions.The GUI supports designing, previewing, and deploying a conversational survey.
- Study platforms: Juji was selected for its Qualtrics-like customization, public accessibility, and richer conversational skills than simple chatbots.These properties were intended to make conversational surveys practical to create and replicate while supporting diverse interaction situations.
- Study scope: The study activates Juji’s available interaction features to assess their holistic effect rather than isolating individual feature contributions.The authors identify controlled feature-level studies as future work because chatbot capabilities are imperfect and open-ended responses are difficult to anticipate.
- Chatbot capabilities: The chatbot can acknowledge input, probe or redirect weak answers, handle excuses and digressions, answer reciprocal questions, and manage side-talking.Its built-in interventions include responses to gibberish, “no opinion,” and reciprocal questions, but the system may still fail to interpret some inputs.
3.3 Survey Questions
The survey used mostly open-ended questions across three parts: warm-up, game-trailer assessment, and additional information. Both chatbot and Qualtrics versions used the same questions in the same order, while trailer order was randomized to reduce bias.
- The survey comprised three parts: warm-up, game-trailer assessment, and additional information.The warm-up included three open-ended questions; the trailer assessment combined open-ended questions with one 1–5 purchase-interest rating.
- Warm up: Participants introduced themselves, discussed favorite games, and identified games they looked forward to playing.
- Game Trailer Assessment: Each participant watched two game trailers and answered questions about reactions, preferences, purchase interest, and buying influences.
- Game Trailer Assessment: All trailer-assessment questions were open-ended except the purchase-interest rating, and trailer order was randomized for each participant.
- Additional Information: Additional questions covered game platforms, information sources, gender, age, and education level.
- The chatbot and Qualtrics surveys used identical question wording and order, supported desktop or mobile completion, and allowed device switching.The chatbot additionally requested optional comments about the survey experience.
3.4 Participants
The study recruited US video gamers aged 18 or older who played at least one hour weekly. A large matching panel pool was randomly divided between Qualtrics and chatbot survey links.
- Participants were US video gamers aged 18 or older who played video games at least one hour per week.
- A panel company queried its large participant database for people matching the study criteria.
- The eligible pool was randomly divided, with one group receiving the Qualtrics link and the other receiving the chatbot link.
3.5 Measures
The study assessed open-ended response quality using manually coded and computational measures guided by Gricean Maxims. Measures covered informativeness, specificity, relevance, clarity, self-disclosure, and response length, with coding reliability checked between researchers.
- The analysis compared chatbot and Qualtrics responses using response-quality and engagement measures.Survey data were stored as question-response pairs, while chatbot-side talking was retained in chat transcripts.
- The quality framework was based on Gricean Maxims and used manual assessment because no effective automatic tool was available.Two researchers independently assessed relevance, specificity, and clarity.
- Informativeness: Informativeness was measured in bits by summing each word’s surprisal, with less frequent words contributing more information.Word frequencies were averaged across four text corpora, and participant-level informativeness aggregated all open-ended responses.
- Specificity: Specificity measured how much concrete detail a response provided, distinguishing shallow, more specific, and detailed descriptions.Examples illustrate specificity levels 0, 1, and 2.
- Self-disclosure was treated as a possible signal of authenticity, while response length excluded participants’ optional comments.
- Relevance: Relevance was rated from 0 to 2 according to whether responses were irrelevant, somewhat relevant, or directly relevant to the question.Participant-level relevance summed the scores across responses.
- The metrics were theoretical estimates rather than unique measures, and alternative metrics or weighting schemes could suit different survey priorities.
- Coder agreement was high, with Krippendorff’s alpha ranging from 0.80 to 0.99 across coding sets.
4 RESULTS
The study compared chatbot and Qualtrics surveys using completion, engagement, and response-quality measures, with analyses controlling for participant characteristics and selected survey-duration factors. Chatbot surveys had higher completion and produced richer, more relevant, specific, and clear responses, although the contribution of individual chatbot behaviors remained unclear.
- Overview: 582 completed surveys comprised 282 chatbot surveys and 300 Qualtrics surveys.Participants ranged from 18–50 years old, and 50% had at least a college degree.
- Analysis: The study compared survey methods with ANCOVA while controlling for demographics, weekly gaming time, and, for selected measures, engagement duration.Response-quality, response-length, and self-disclosure analyses additionally controlled for engagement duration.
- Response quality: Chatbot surveys collected 39% more information and 12% more relevant responses than Qualtrics surveys.The survey method significantly contributed to both differences after the stated controls.
- Response quality: Chatbot surveys produced 25.7% better overall response quality, with responses that were more relevant, specific, and clear.The overall response-quality index aggregated relevance, clarity, and specificity; participants with completely irrelevant responses were also less common in the chatbot condition.
4.3 RQ2: How Would a Chatbot Impact Participant Engagement?
Compared with Qualtrics, the chatbot increased engagement across completion time, response length, and self-disclosure. Participants also generally reported positive experiences with the conversational survey.
- Engagement duration: Seven more minutes were spent on average completing chatbot surveys than Qualtrics surveys, with the duration difference statistically significant.Survey method was the only significant factor after controlling for demographics, game-playing time, and response length.
- Engagement duration: Chatbot survey completion was 54%, compared with 24% for Qualtrics, despite participants being paid only for completing the survey.The authors interpret this combination as indicating willingness to remain engaged with the chatbot longer.
- Response length: Participants contributed 30 more words on average in chatbot surveys than in Qualtrics surveys.This difference remained after controlling for demographics, game-playing time, and time spent with the chatbot.
- Self-disclosure: Participants disclosed 1.6 more types of personal information on average in chatbot surveys, including personal information from 32.62% versus 15.67%.The chatbot condition elicited details about personal facts, daily activities, and personality, whereas Qualtrics responses were mostly about preferred game types.
- Self-disclosure: Reciprocity may have contributed to greater self-disclosure after the chatbot greeted participants and introduced itself before asking for a self-introduction.This is presented as a conjecture based on prior human-agent interaction research.
- Participants’ feedback: Among 282 completed chatbot surveys, 70% left optional comments, and 95% of those comments were positive.Positive comments commonly expressed personal connection or enjoyment; about a quarter called the chatbot survey fun or their best survey experience.
4.4 Summary of Findings
The study found that chatbot surveys produced higher-quality open-ended responses, greater participant engagement, and generally favorable participant reactions than Qualtrics surveys.
- Response quality: Chatbot participants provided significantly more relevant, specific, and clear responses to open-ended questions than Qualtrics participants.The finding concerns response quality across the reported dimensions rather than a single quality measure.
- Participant engagement: Participants engaged more with the chatbot by spending more time, writing longer responses, and disclosing more information in greater depth and scope.The engagement finding combines duration, response length, and self-disclosure measures.
- Participant reactions: Participants’ comments indicated that a majority enjoyed chatting with the chatbot and preferred this conversational survey format for future surveys.The authors note that novelty and positivity toward humanized machines may have influenced these comments.
5 DISCUSSION
The discussion connects richer responses and engagement with conversational interaction while examining positivity bias, survey fatigue, privacy, and limits on generalization and causal attribution.
- Quality and positivity bias: Chatbot surveys elicited richer, deeper responses without causing unwanted positivity bias in trailer ratings.The authors report that chatbot use did not significantly increase positive ratings.
- Survey fatigue: Participants spent over 20 minutes with the chatbot, and interactive conversation appeared to help overcome survey-taking fatigue.The authors support this interpretation with enjoyment comments from 42.2% of chatbot participants.
- Inferred characteristics: A chatbot may elicit participant information while simultaneously inferring participant characteristics, potentially reducing the need for separate surveys.The authors describe this as a preliminary potential benefit based on relationships between inferred gamer characteristics and purchase interest.
- User privacy and control: Privacy controls can protect users and improve engagement, but obfuscating information may interfere with collecting authentic data.The paper calls for future work balancing privacy protection with information validity.
- Study audience and scope: The gamer-focused sample may limit applying the findings to other populations, particularly because gamers may be more receptive to chatbots.The authors also report that game-playing time contributed to observed differences.
- Novelty effects: Novelty from the conversational format and the chatbot’s rich skills may have affected participant behavior, but the study could not control for this effect.The authors expect some novelty effects may diminish as conversational surveys become more common.
6 CONCLUSIONS
The field study found that AI-powered chatbot surveys elicited richer free-text responses and more participant engagement than a form-based online survey. The findings suggest chatbot surveys as a promising approach for collecting open-ended responses and addressing survey-taking fatigue.
- About 600 participants completed either a Juji chatbot survey or a Qualtrics form-based survey, with over 5200 free-text responses analyzed.
- Participants using the chatbot produced significantly more relevant, specific, and clear free-text responses than Qualtrics participants.
- Chatbot participants spent more time, wrote longer responses, and disclosed more information about themselves.
- 190 participants, or 67.4%, reported a positive experience and willingness to take surveys in a chat format.
- The study suggests chatbot surveys as a promising method for collecting open-ended responses and overcoming survey-taking fatigue.
- The results provide design implications for creating and employing chatbots for survey success.