Source-linked AI summary
The State of Speech in HCI: Trends, Themes and Challenges
Leigh Clark, Phillip Doyle, Diego Garaialde, Emer Gilmartin, Stephan Schlögl, Jens Edlund, Matthew Aylett, João Cabral, Cosmin Munteanu, Benjamin Cowan
TL;DR
Speech interfaces are increasingly common, but the core trends, themes, and methods of user-focused speech HCI research were unclear. The paper reviews 68 empirical studies from core HCI venues and maps their systems, measures, topics, and challenges. It finds predominantly usability- or theory-focused research using Wizard of Oz systems, prototypes, commercial systems, and self-report questionnaires, alongside needs for stronger theory, design work, broader contexts, and more reliable measures.
Problem
User-focused speech HCI research was limited and lacked a clear account of the field’s core trends, topics, and methods.
Method
The authors reviewed 68 empirical speech HCI papers selected through adapted PRISMA procedures and analyzed their aims, interaction types, methods, topics, interfaces, measures, and participants.
Results
Most research was usability- or theory-focused or examined wider system experiences, commonly using Wizard of Oz systems, prototypes, commercial systems, and self-report questionnaires.
Takeaways & Limitations
The review identifies needs for stronger interaction theory, more design work, evaluation across multiple-user contexts, improved measurement reliability, validity and consistency, in-the-wild deployment, and lower technical barriers.
Takeaways & Limitations
The review excluded relevant work from other fields, non-empirical research, and embodied conversational agents and robotics.
Abstract
from arXiv · showhide
Speech interfaces are growing in popularity. Through a review of 68 research papers this work maps the trends, themes, findings and methods of empirical research on speech interfaces in HCI. We find that most studies are usability/theory-focused or explore wider system experiences, evaluating Wizard of Oz, prototypes, or developed systems by using self-report questionnaires to measure concepts like usability and user attitudes. A thematic analysis of the research found that speech HCI work focuses on nine key topics: system speech production, modality comparison, user speech production, assistive technology \& accessibility, design insight, experiences with interactive voice response (IVR) systems, using speech technology for development, people's experiences with intelligent personal assistants (IPAs) and how user memory affects speech interface interaction. From these insights we identify gaps and challenges in speech research, notably the need to develop theories of speech interface interaction, grow critical mass in this domain, increase design work, and expand research from single to multiple user interaction contexts so as to reflect current use contexts. We also highlight the need to improve measure reliability, validity and consistency, in the wild deployment and reduce barriers to building fully functional speech interfaces for research.
1. Introduction
Speech interfaces are increasingly prominent, but the state of user-focused speech HCI research remains unclear. This review maps empirical work in the field to identify its methods, themes, and challenges.
- Speech is becoming a prominent way to interact with automatic systems, including telephony, IVR interfaces, IPAs, and home-based devices.The review situates speech HCI within expanding use across devices and interaction settings.
- User-side speech interface research has been described as limited compared with speech technology research.
- Most reviewed research is usability- or theory-focused, or examines wider system experiences.
- The research commonly uses Wizard of Oz systems, prototypes, or established commercial systems, alongside self-report measures of usability and user attitudes.
2. What are speech interfaces?
Speech interfaces use speech as system output, user input, or two-way dialogue. The reviewed work frames these interactions across systems ranging from monologic interfaces to spoken dialogue systems.
- Speech interfaces use prerecorded or synthesised speech to communicate or interact with users.
- Speech can serve as primary input, primary output, or part of a two-way dialogue.
- Monologic interfaces communicate in one direction, with speech produced only by the system or supplied only by the user.
- Spoken dialogue systems support spoken natural-language interaction, from formulaic question-answer exchanges to systems allowing broader queries, interruption, and task changes.The passage notes that current system dialogues remain quite stiff.
- The research aims to map speech-based HCI because the field lacks a clear account of its core trends, topics, and methods.
3. Method
The review identified empirical speech HCI studies through database searches and explicit selection criteria. Independent screening produced a final set of 68 papers for analysis.
- 3.1. Scope: The review covered 68 publications on speech as system output, user input, or user-system dialogue.
- 3.2. Search procedure: Searches used ACM Digital Library, ProQuest, and Scopus, with terms drawn from speech literature and a survey of 11 leading researchers.
- 3.3. Inclusion and exclusion criteria: The search produced 1,181 unique entries after duplicate removal before inclusion and exclusion criteria were applied.
- 3.3. Inclusion and exclusion criteria: Embodied interfaces and studies without empirical measurement were excluded to avoid embodiment-related confounds and retain user-evaluated research.
- 3.3. Inclusion and exclusion criteria: Independent filtering by two authors showed strong agreement, with Kappa = 0.843 (p < .001), before 68 papers were selected for analysis.
4. Results: Publication Trends & Research Methodologies
Speech HCI research was concentrated in conference publications and computer-based or telephone contexts, using mock, prototype, or existing systems. Studies most often combined objective and subjective measures, especially questionnaires assessing attitudes and usability.
- 4.1. Publication Trends: 63.2% of reviewed papers (N=43) were published in conferences, compared with 36.8% (N=25) in journals.CHI accounted for 39.7% (N=27) of all papers, while the International Journal of Human-Computer Studies accounted for 14.7% (N=10).
- 4.2. Research Methods: One third of systems (N=26) were Wizard of Oz mock systems, while 45 papers focused on existing systems or working prototypes.Wizard of Oz studies concealed a confederate controlling system output to assess responses to hypothetical systems.
- 4.2.2. Device contexts: 48.6% of systems (N=35) used computer-based interactions, and 33.3% (N=24) explored telephone-based dialogues.Six systems were mobile applications or smartphone IPAs, and four were vehicle-based systems.
- 4.2.4. Measures: The majority of papers (N=37) combined objective and subjective measures; 19 used only subjective measures and 12 used only objective measures.
- 4.2.4. Measures: User attitudes, task performance, lexis and syntax, and perceived usability were recurring measurement concepts.Questionnaires were especially common, with 40 papers using them and 38 relying mainly on custom-built scales.
- 4.2.5. Data collection methods: Observations were common (N=36), interviews were less common (N=20), and questionnaires were used in 40 papers.
5. Results: Research Topics
The review categorizes the included papers by type of work and identifies the primary research themes covered across them.
- The review categorizes the types of work represented in the papers.
- It also categorizes the primary research themes covered across the reviewed papers.
- These categories organize the review’s analysis of speech HCI research.
5.1. Type of work
The reviewed papers divided into usability/theory-based research and system experience research. Usability/theory-based work was the larger category, while system experience research examined interactions with working prototypes or existing systems.
- 43 papers were categorized as usability/theory-based research, examining design choices, behaviors, usability, performance, user experience, or concepts borrowed from human communication.Examples included alignment and vague language, alongside modifications to specific components such as synthesized speech.
- 25 papers were classified as system experience research, exploring user interactions with working prototypes or existing systems rather than specific components or theory-based concepts.This category included exploratory studies using semi-structured interviews, ethnography, and other qualitative techniques to identify user views and issues.
5.2. Research Themes
The thematic analysis organized the reviewed speech HCI research around system speech production, modality comparison, and related interaction topics. System speech production was the largest topic, while modality-comparison findings were mixed across input and output combinations.
- Research themes: Thematic coding was conducted independently by two authors, with inconsistencies resolved through discussion involving an independent observer.
- System speech production: 18 papers, or 26.5%, focused on system speech production and its effects on interaction behavior or user experience, making it the largest observed category.The category covered system-directed speech and was divided into four sub-topics.
- System speech production: Studies of synthesized speech examined voice characteristics such as personality, accent matching, and voice matching, with effects on perceived social presence, informativeness, likeability, comprehension, and driving performance.Matching email voices to sender characteristics improved comprehension but negatively affected driving performance; one medical IVR study found no effects of user gender or voice personality on behavior.
- System speech production: Speech rate and message timing also shaped outcomes: fast speech was more arousing but less understandable, while proactive road-condition alerts lowered collisions and turning errors.Younger participants preferred fast news delivery, whereas older participants preferred slower delivery; more descriptive alerts were perceived more positively than shorter general messages.
- Modality comparison: 16 papers compared modalities, including speech, keyboard/text, mouse, gesture, graphics, and pen input, with results that varied by task and combination.Participants sometimes preferred speech, sometimes traditional input, and sometimes multimodal interaction for specific activities.
- Modality comparison: Multimodal preferences were task-dependent: users preferred mouse navigation and natural-language data entry, multimodal correction, and speech when their hands were occupied.Speech use increased when graphical input became less efficient, and spoken feedback increased perceived human-likeness in one timetable system.
5.2.3. User speech production
Six reviewed papers examined how users produce speech in interactions with computers and other partners. They addressed language adaptation, addressee identification, and alignment with conversational partners.
- Six papers focused on user speech production, including studies of general speech behavior, addressee identification, and alignment.
- General user speech production: People adapted language choices according to their partner models, while differences between human- and computer-directed speech choices decreased with interaction familiarity.In computer interactions, people also tended to use fewer fillers, request confirmation and repetition more often, and make fewer topic shifts.
- Addressee identification: Users indicated self-talk with lower vocal loudness than system-directed speech, supporting addressee identification through amplitude differences.
- Alignment: Although synthesis humanness affected people’s partner models, it did not significantly change syntactic alignment levels between human and computer partners.
- Assistive technology and accessibility: Six papers addressed assistive technology and accessibility, with most deploying speech interfaces for people with specific needs or requirements.Assistive technology was framed as adapted or specially designed equipment, products, and technologies that assist daily living.
5.2.5. Design Insight
The reviewed studies used speech-interface evaluations to derive design insights across IVRs, development contexts, IPAs, memory, and engagement. Findings show that dialogue structure, accessibility, user expectations, and evaluation perspective shape interaction outcomes.
- Five studies generated design insight for subsequent speech-interface development, including guidelines for conversational structures, lexicons, and information flow.
- Experiences with IVR systems: IVR studies found that dialogue options, confirmation strategies, menu visibility, and adaptation to ASR problems affected efficiency, task completion, and usability.Older and younger users were more efficient when scheduling systems presented more options and avoided explicit confirmations; hidden banking-menu options could prevent successful completion, while ASR-aware conservative dialogue improved train-schedule retrieval.
- Speech technologies for development: Four studies examined speech technologies supporting rural communities, novice users, or people with low literacy, comparing speech with text and rich multimedia interfaces.
- People’s experiences with IPAs: IPA research identified mismatches between users’ mental models and system capabilities, alongside limited trust and barriers including usability difficulties, public-use embarrassment, and negative effects of human likeness.
- User memory: Memory studies found impaired recall with five or more IVR options and suffixes, while working-memory span did not affect appointment recall in a healthcare-scheduling study.
- Engagement evaluation: Crowd evaluations produced consistent engagement ratings across human and computer communication conditions, but self-evaluations correlated weakly with third-party ratings.
6. Discussion
The review maps a fragmented speech-HCI field whose empirical work is dominated by usability and theory-focused studies, with recurring weaknesses in theory coverage, design practice, measurement, context, and research accessibility. It identifies priorities for building a stronger evidence base that better reflects contemporary speech use.
- Speech-HCI research coalesces around nine topics, but most studies are usability/theory-focused, quantitative, and based on prototypes, Wizard of Oz systems, or commercial systems.
- Theory: Existing theories explain user attitudes toward speech design, but comparable theories for users’ language choices remain lacking.
- Achieving critical mass: Most topics contain fewer than seven papers and fragmented subtopics, making consensus and synthesis of results difficult.
- Design work: Few reviewed studies used interaction-design approaches, leaving no clear design considerations or robust heuristics for user-centred speech interactions.
- Evaluation measures: Self-report measures often lacked reliability or validity testing, and inconsistent concepts across studies hinder development of robust measures and cumulative knowledge.
- Research contexts: Research relied heavily on lab-based prototypes and Wizard of Oz systems, requiring further work on transfer to real-world settings and comparisons across contexts.
- Research accessibility: The scarcity of design research may reflect perceived barriers to building operational prototypes, although packages, toolkits, and platform kits exist to support prototyping.
- Scope: The review’s core-HCI and empirical-study scope excluded relevant work from other fields and research without user evaluation or user-based data collection.
7. Conclusion
The paper maps speech-interface research in HCI through a review of 68 papers, identifying nine topics, common methods, and challenges for future theory, design, evaluation, context, and implementation work.
- A review of 68 core-HCI papers identified nine primary research topics and methodological, research, and evaluation challenges for speech-interface research.