Source-linked AI summary
Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review
Rock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas, Ziang Xiao, Emily Tseng, Danielle Bragg
TL;DR
HCI lacks a clear account of how LLMs are being adopted and studied. The paper systematically reviews 153 CHI papers from 2020–2024, finding diverse applications, five LLM roles, and recurring validity and reproducibility concerns. It uses these findings to propose research opportunities and guiding questions for more rigorous, appropriate LLM-related work.
Problem
There has been limited understanding of LLM uptake in HCI despite their potential to reshape interfaces, sociotechnical systems, and research practices.
Method
The authors conduct a systematic literature review of 153 LLM-related CHI papers from 2020–2024 and taxonomize applications, roles, contributions, limitations, and risks.
Results
The review finds LLM work across 10 domains, primarily through empirical and artifact contributions, with five roles and recurring validity and reproducibility concerns.
Takeaways & Limitations
The paper offers research opportunities and guiding questions for evaluating task appropriateness, validity, reproducibility, and consequences throughout LLM-powered HCI projects.
Takeaways & Limitations
The manual and iterative review process limited the authors’ ability to conduct a more expansive literature review.
Abstract
from arXiv · showhide
Large language models (LLMs) have been positioned to revolutionize HCI, by reshaping not only the interfaces, design patterns, and sociotechnical systems that we study, but also the research practices we use. To-date, however, there has been little understanding of LLMs' uptake in HCI. We address this gap via a systematic literature review of 153 CHI papers from 2020-24 that engage with LLMs. We taxonomize: (1) domains where LLMs are applied; (2) roles of LLMs in HCI projects; (3) contribution types; and (4) acknowledged limitations and risks. We find LLM work in 10 diverse domains, primarily via empirical and artifact contributions. Authors use LLMs in five distinct roles, including as research tools or simulated users. Still, authors often raise validity and reproducibility concerns, and overwhelmingly study closed models. We outline opportunities to improve HCI research with and on LLMs, and provide guiding questions for researchers to consider the validity and appropriateness of LLM-related work.
1 Introduction
LLMs are rapidly reshaping HCI research and raising questions about how HCI methods influence LLM development and societal outcomes. This review examines their uptake in CHI and offers guidance for rigorous, responsible research.
- LLMs are being used across the HCI research pipeline, from ideation and system development to data analysis and paper-writing.
- The growth of LLM research has prompted HCI researchers to debate opportunities, challenges, and responsible practices through studies, workshops, and social-media discussions.
- HCI methodologies increasingly shape LLM development through techniques such as reinforcement learning from human feedback, while other communities scrutinize potential societal harms.
- The review analyzes 153 CHI papers from 2020–2024 to identify application domains, LLM roles, contribution types, and author-articulated concerns.
- The authors identify 10 application domains, five LLM roles, predominantly empirical and artifact contributions, and 29 reported limitations.
- The paper provides future research opportunities and guiding questions addressing the validity and appropriateness of LLM-related studies.
2 Related Work
Prior literature reviews establish systematic review as a way to identify patterns and limitations, while recent work documents rapid LLM growth across disciplines. This study extends that work with a qualitative analysis focused on CHI.
- Systematic literature reviews help HCI identify patterns, trends, limitations, and shared conceptual frameworks.
- Existing HCI reviews have qualitatively analyzed large samples to characterize contribution types, communities, methods, topics, and research areas.
- Reviews outside HCI have examined LLM training, evaluation, retrieval-augmented generation, multi-agent systems, and societal implications.
- A large arXiv analysis found society-facing and HCI topics were among the fastest-growing areas of LLM research.
- This work focuses on CHI papers to examine where authors apply LLMs and how they leverage them to make HCI contributions.
- The study extends quantitative trend analysis with an in-depth qualitative analysis of recent CHI literature.
- The authors identify LLM roles and reported limitations in HCI while advocating proactive attention to limitations and research rigor.
3 Methods
The study reviews generative-LLM research in CHI proceedings from 2020–2024 using iterative human coding and an adapted PRISMA-based sampling process. The resulting corpus is deliberately generative rather than exhaustive.
- The review examines generative LLMs in CHI from 2020–2024, coding contribution types, LLM roles, and disclosed limitations.
- CHI was selected as a flagship, rigorously peer-reviewed venue spanning diverse HCI application areas and methodologies.
- The sample is considered generative rather than exhaustive because it covers CHI rather than all SIGCHI conferences.
- The authors assembled CHI full-text proceedings and applied an adapted PRISMA process to filter papers for LLM relevance.
- The final corpus contained 153 papers after keyword filtering and false-negative validation found one additional relevant paper among 200 initially excluded papers.
- The search excluded full-text matching and broad artificial-intelligence terms to reduce false positives and focus specifically on LLM-related work.
- An iterative codebook process used existing taxonomies, repeated independent coding, consensus-based refinement, and Krippendorff’s alpha to guide reliability discussions.
4 Results
The analysis maps where LLMs are applied in CHI, how researchers use them, what contributions they make, and which limitations and risks authors report.
- The review taxonomizes LLM application areas, research uses, HCI contributions, and author-articulated limitations and risks.
- Its results address both how HCI applies LLMs and how HCI studies them.
- The limitation taxonomy captures recurring concerns reported by authors across the reviewed papers.
4.1 Application Domains
CHI researchers applied LLMs across 10 diverse domains, with communication and writing the most-studied area. Applications span productivity, education, responsible computing, programming, health, design, accessibility, creativity, and reliability.
- Communication and writing: Communication and writing was the most-studied domain, covering writing tasks and AI-mediated communication.It accounted for 22.88% of papers (N=35).
- Programming: Programming research included code generation, no-code platforms, code explanation, programming education, and prompt engineering.One analysis found that 52% of ChatGPT answers to 517 StackOverflow questions contained incorrect information and 77% were verbose.
- Additional application areas: Other domains included reliability and validity, health, design, accessibility and aging, and creativity.Reliability and validity accounted for 10.46% (N=16), well-being and health 9.15% (N=14), design 8.50% (N=13), accessibility and aging 7.84% (N=12), and creativity 5.88% (N=9).
- Reliability and validity: Reliability-oriented work evaluated LLM validity and built tools for hallucination detection, interactive evaluation, and more transparent prompt construction.These tools aimed to improve users’ caution, evaluation control, task outcomes, transparency, controllability, and collaboration with black-box models.
4.2 Contribution Types
LLM-focused CHI papers were primarily empirical and artifact-oriented, often combining tool construction with user evaluation. Methodological, theoretical, and dataset contributions appeared less frequently.
- Empirical and artifact contributions: 98.70% of papers made empirical contributions, while 61.44% made artifact contributions involving a tool.These contribution types frequently appeared together when authors built an artifact and empirically tested it with users.
- Artifact contributions: Artifact contributions ranged from open-source systems to wireframes, with LLMs varying from pipeline-wide use to textual-data processing.The authors applied the artifact code when papers claimed that LLMs were or would be part of the system.
- Less frequent contribution types: The sample contained 16 methodological contributions, 8 theoretical contributions, and 6 dataset contributions.Methodological examples included UX evaluation, synthetic user data, and creativity metrics; theoretical examples included frameworks and design spaces.
- Less frequent contribution types: The review found one survey contribution and no opinion contributions, while curating real-user benchmark datasets at scale remained challenging.Synthetic datasets may lower barriers to large, diverse evaluations, but the authors distinguish this from curating real-user datasets.
4.3 LLM Roles
The review identifies five roles for LLMs in HCI research, spanning system building, research support, simulated participation, and studying LLMs themselves. Most sampled projects use LLMs as system engines, while newer methodological roles commonly include experimental validation.
- Scope of the taxonomy: The five-role taxonomy reflects the sample, which primarily offers empirical contributions, but may not fit every interdisciplinary HCI project.Figure 3 maps these roles across common HCI research stages and empirical study types.
- LLM roles: 62.74% of papers used LLMs as system engines within systems, prototypes, algorithms, or programming frameworks.These systems can generate ideas, code, or conversations.
- LLM roles: LLMs served as research tools for data collection, analysis, or writing, including GPT-4-assisted qualitative coding after manual codebook development.All 15 papers in this role justified the usage, and all but one provided further experimental validation.
- LLM roles: LLMs generated synthetic research data, including artificial greeting messages and natural-language datasets, which authors then analyzed as part of their contributions.These studies used LLMs to create data for research purposes rather than only to support analysis.
- LLM roles: LLMs acted as participants or users by simulating human responses, personas, user feedback, or usability feedback.10 of 11 papers provided textual justification and experimental validation for this methodology.
- LLM roles: LLMs were also studied as objects through their training datasets, response outputs, and properties such as hallucination.This role examines problems and mechanisms inherent to LLMs themselves.
4.4 Limitations
Authors report limitations involving LLM performance, resources, evaluation, and research validity. These concerns include representational bias, incomplete training coverage, nondeterministic or fabricated outputs, opaque errors, changing models, unreleased prompts, and constrained reproducibility.
- LLM performance: 11.11% of papers reported LLM bias toward different groups, including stereotyped representation and failures to model some user groups.Examples include outputs favoring white men and chatbots failing to recognize nuanced LGBTQ+ identities and experiences.
- LLM performance: 9.80% of papers cited limited or outdated training-data coverage, including GPT-4 underperformance in non-English languages.Uncertainty about whether study data appeared in model training also constrained interpretation.
- LLM performance: 7.84% of papers reported nondeterministic responses that could change unpredictably for the same prompt.Authors noted that temperature-zero sampling or guided generation can alleviate this problem.
- LLM performance: 8.50% of papers reported hallucination, meaning LLMs can produce inaccurate or fabricated information.RAG may help, but applications using it can still suffer hallucination issues.
- LLM performance: 16.99% of papers described unspecified errors and biases, often because models are opaque and their outputs are difficult to control or replicate.Authors sometimes observed inaccurate outputs without identifying the precise causes.
- Resource limitations: Resource limitations included computational cost, financial cost, and restricted token windows that constrained local execution, reproducibility, or input size.Using open-source models could require substantial GPU resources, while API use and subscriptions increased monetary costs.
- Evaluation limitations: 16.99% of papers cited a lack of evaluation standards or metrics for judging correctness, safety, conversational context, or domain-specific outputs.Authors described benchmarks and evaluation methods as active or open research areas.
- Research validity: Validity concerns included small or non-diverse samples, reliance on closed GPT-family models, model changes over time, and unreleased prompts.Of 153 papers, 130 used or studied closed GPT-family models; among 146 prompting studies, 40.4% did not release prompts.
5 Discussion
LLM research at CHI has grown substantially, prompting examination of where the community is focusing and what this surge means for HCI’s prototyping, design, and research norms. The authors respond with guiding questions centered on task appropriateness, validity, reproducibility, and consequences.
- Research growth: CHI has experienced substantial growth in research studying LLMs, echoing trends in other fields.The discussion examines the community’s focus and the implications for HCI norms around prototyping and design.
- Guiding questions: The paper proposes guiding questions for HCI researchers to assess task appropriateness, validity, reproducibility, and consequences throughout LLM-powered projects.The proposal is intended to support reflection on research rigor.
5.1 Revealed Growth Opportunities for HCI
The review identifies broad opportunities for HCI to deepen and diversify LLM research, beyond the field’s current concentration on artifacts and empirical evaluations. It highlights underexplored domains, contribution types, benchmarking, theory, methods, and prototyping standards.
- Application domains: LLMs span diverse HCI application domains, with Communication & Writing receiving the most attention.The authors describe this concentration as potentially related to LLMs’ direct relevance to language production.
- Contribution types: LLM-related HCI papers predominantly make artifact and empirical contributions, often through user studies of new artifacts.The authors encourage greater attention to five other contribution types that are less represented.
- Datasets and benchmarking: Community-driven, participatory benchmarking could produce datasets representing diverse user requirements and downstream harms.The authors position HCI’s sociotechnical orientation as well suited to innovating benchmarking practices.
- Theory and methods: The field needs more theoretical, methodological, literature-review, and opinion contributions to develop transferable principles and shape understanding of LLMs.The authors specifically point to design-space work across application domains.
- Prototyping standards: LLM proliferation raises questions about prototype fidelity, artifact contributions, system contributions, and the continuing value of Wizard of Oz approaches.These questions affect prototyping standards, peer-review norms, research validity, and valued contributions.
5.2 Challenges: Validity, Reproducibility, and Consequences
The review finds that LLM research faces substantial validity and reproducibility challenges, alongside limited reporting of broader consequences. Closed-model dependence and underspecified errors make results difficult to interpret and reproduce.
- Validity and reproducibility: LLM research is growing rapidly despite authors’ frequent concerns about research validity, increasing urgency around reproducibility.The authors connect this trend to broader calls for examining reproducibility in HCI.
- Validity and reproducibility: 84.98% of sampled papers studied closed GPT-family models, creating reproducibility and interpretability concerns.The sample included 61 GPT-4, 41 GPT-3.5, and 26 GPT-3 papers; closed models may change over time and conceal prompts, training data, or weights.
- Validity and reproducibility: Authors commonly noted limited training data, hallucination, and nondeterministic responses as threats to expected behavior or external validity.These concerns applied whether LLMs were studied directly or used to power an interactive system.
- Validity and reproducibility: 16% of papers reported unspecified errors and biases, indicating that their precise performance effects often remained unaddressed.The authors argue that specifying the nature of errors or bias is critical for understanding effects on systems and results.
- Consequences and risks: Only 35 papers discussed potential consequences, often in an ethics statement, despite widespread attention to validity and reproducibility.The authors distinguish limitations affecting conclusions from consequences involving long-term social impact and deployment.
- Consequences and risks: Structured consideration of consequences could help HCI establish norms for rigorous and thoughtful LLM research as deployment expands.The authors propose ethics statements or other structured mechanisms for reflecting on consequences.
5.3 Guiding questions for HCI researchers using LLMs
The authors propose iterative, open-ended questions for evaluating whether and how LLMs should be used throughout HCI research. The guidance addresses roles, model choice, disclosure, and role-specific limitations.
- Role-specific reflection: Researchers should examine limitations associated with each selected LLM role rather than treating LLM use as categorically harmful or appropriate.The guidance is intended as iterative reflection rather than a completion checklist.
- Project decisions: Researchers should identify the LLM’s project role and consider whether simpler models or humans could achieve the same result.The guidance treats whether an LLM is needed at all as a central question.
- Project decisions: Model selection should weigh transparency, reproducibility, ethics, and use-case fit when choosing between open and closed models.Closed models may still be appropriate for audits, cheap rapid prototypes, or other cases described by the authors.
- Documentation: Researchers should disclose exact model versions, full prompts, or user-facing prompt templates.The authors note that names such as gpt-4o-2024-08-06 and gpt-4o can refer to different models.
- LLMs as system engines: For LLM-powered systems, researchers should match artifact fidelity to the contribution and assess how models and prompts affect performance.For formative studies or user perceptions of outputs, a Wizard of Oz approach may be more appropriate than a fully deployed commercial API system.
- LLMs as objects of study: When studying LLMs directly, researchers should define the behaviors of interest and evaluate outputs against real discourse and deliberation.The authors warn that human language can inadequately capture behavioral fidelity, threatening validity.
5.4 Limitations
The review’s conclusions are constrained by potentially incomplete sampling and a labor-intensive manual qualitative process. These limits restricted the review’s scope and leave room for broader follow-up studies.
- Sampling: The sampling approach may have missed papers that used LLMs without the review’s selected keywords.A robustness check found one paper mentioning GPT-4 only once in its methods.
- Future scope: Future research could examine LLMs’ impact in other HCI subcommunities and conferences beyond CHI.The authors also identify challenges in sharing datasets for reproducing fine-tuning and multi-agent configurations.
- Review process: Manual and iterative review work limited the scope of the literature review.The authors spent substantial effort curating papers, resolving coding disagreements, and discussing difficult cases.