Source-linked AI summary

Evaluating Human and LLM-Generated Thematic Analysis in HRI for Vulnerable Populations: A Comparative and Ethical Analysis

Alva Markelius, Fethiye Irmak Dogan, Julie Bailey, Hatice Gunes

arXiv:2608.21420v1cs.ROcs.HC

TL;DR

LLM-generated thematic analysis remains insufficiently examined in sensitive HRI research involving vulnerable populations, where interpretive validity and ethics matter. The paper compares human and LLM analyses across thematic grouping, semantic similarity, and ethical patterns in disagreement, finding partial but uneven alignment with systematic interpretive differences. It concludes that LLMs may support early-stage analysis, while human oversight remains important for higher-level interpretations involving identity, power, and lived experience.

  • Problem

    Evidence on whether LLM-generated thematic analysis generalises appropriately to HRI research with vulnerable populations is limited, particularly regarding ethical and systematic interpretive differences.

  • Method

    The study compares human- and LLM-generated thematic analyses of interview data from 31 university students with disabilities, evaluating thematic grouping, semantic similarity, and disagreement patterns.

  • Results

    Human and LLM analyses show partial alignment: grouping performs above chance and semantic similarity is moderate, while disagreements systematically shift interpretation, abstraction, and representation.

  • Takeaways & Limitations

    LLMs can support early-stage coding and clustering, but human-in-the-loop review remains valuable for higher-level interpretations involving identity, power, and lived experience.

  • Takeaways & Limitations

    The analysis includes several cases in which LLMs replace participants’ specific terms with broader labels, raising concerns about semantic sanitisation and representational fidelity.

Abstract

from arXiv · show

Thematic analysis (TA) has long been regarded as an inherently human, reflexive, and interpretive process. However, the extent to which LLM-generated TA is appropriate for Human-Robot Interaction (HRI) research involving vulnerable populations remains largely unexamined and raises critical questions about validity and ethics, particularly in sensitive research contexts. This paper presents a comparative study of human- and LLM-generated TA in an HRI context with a focus on vulnerable populations. We evaluate both objective and semantic agreement between human- and LLMgenerated themes, and examine whether observed divergences reflect systematic interpretive patterns with ethical significance. Our analysis investigates whether LLM-generated TA risks marginalising or misrepresenting the experiences of vulnerable participants, with implications for researchers employing LLM-assisted TA in HRI.

I. INTRODUCTION

The paper examines whether LLM-generated thematic analysis is appropriate for sensitive HRI research involving vulnerable populations. It compares human and LLM analyses while investigating agreement, semantic alignment, and ethically significant interpretive divergences.

  • Thematic analysis is an interpretive craft shaped by contextual judgment, reflexivity, subjectivity, domain knowledge, cultural awareness, and ethical sensibility.
  • LLMs have shown promising agreement with human analysts, but their generalisability to HRI research involving vulnerable populations remains largely unexamined.
  • The study compares human- and LLM-generated thematic analyses of interview data from 31 university students with disabilities in an HRI application.
  • The analysis asks whether human and LLM annotations share thematic groupings and semantically similar labels at code, subtheme, and theme levels.
  • It also examines whether disagreements are systematic and risk marginalising or misrepresenting vulnerable or marginalised participants’ experiences.

II. RELATED WORK

Prior work reports efficiency and agreement in LLM-assisted thematic analysis, but leaves important methodological, epistemological, and ethical concerns unresolved. This paper addresses limited attention to systematic interpretive differences in vulnerable-population research.

  • LLM-assisted thematic analysis can improve efficiency in open coding and theme generation, but may introduce contextual oversimplification, bias amplification, and premature interpretive closure.
  • Reported challenges include hallucination, model-version variability, unstable excerpt extraction, and the interpretive labour needed for supervision and ethical governance.
  • Existing research has focused mainly on general qualitative datasets, with less attention to ethical implications in studies involving vulnerable populations.
  • Prior evaluations have predominantly measured aggregate agreement rather than deeper epistemic or normative divergences between human and LLM analyses.

B. Thematic Analysis in HRI for Vulnerable Populations

Thematic analysis is established in HRI research with vulnerable populations as a way to capture lived experiences and support participant authority. The study evaluates whether LLM-assisted analysis preserves that role in a disability-focused HRI dataset.

  • TA is widely used in HRI studies involving vulnerable and marginalised communities, including children, older adults, people with chronic illness, and indigenous learners.
  • In vulnerable-population HRI, TA can support participants’ authority to challenge the norms, perceptions, and biases embedded in the field.
  • The study evaluates both consistency with human analysis and whether divergences carry ethical weight when analysing data from vulnerable populations.
  • The dataset comes from a prior within-subjects HRI study of social robots and voice agents as mediation-support tools for disabled university students.

A. Human Thematic Analysis

The study establishes human and LLM thematic-analysis procedures using Braun and Clarke’s six-phase framework. Human analysis relied on researcher expertise and corpus immersion, while LLM analysis used sequential prompting and cross-batch normalisation.

  • Human Thematic Analysis: A researcher with expertise in disability and neurodivergence conducted reflexive thematic analysis using Braun and Clarke’s six-phase framework.
  • Human Thematic Analysis: Human familiarisation involved active engagement with the corpus during transcription editing and video review before salient segments were extracted for analysis.
  • Human Thematic Analysis: The LLM analysis adapted prior prompting strategies and assigned each phase of Braun and Clarke’s framework as a separate sequential task.
  • Human Thematic Analysis: Pooled reanalysis normalised outputs across batches by collapsing redundant codes for similar phenomena into shared themes.

C. Clustering Analysis

The study evaluates clustering consistency by comparing human and LLM assignments of utterances and semantic alignment by comparing corresponding labels with cosine similarity.

  • Clustering consistency: Adjusted Mutual Information compares human and LLM clusterings while correcting agreement for chance.The analysis applies AMI separately at theme and subtheme levels.
  • Clustering consistency: AMI values indicate agreement relative to chance: 1 is perfect agreement, 0 is chance-level agreement, and negative values indicate below-chance agreement.
  • Semantic alignment: Human and LLM labels are embedded into vectors, then compared using cosine similarity for codes, subthemes, and themes.The analysis reports mean similarity and standard deviation across utterances within each human theme.

E. Qualitative Analysis

The qualitative analysis examines selected high- and low-similarity codes to compare human and LLM interpretations and assess representational differences.

  • Three highest- and three lowest-similarity codes were selected for qualitative comparison.High-similarity cases provide a baseline, while low-similarity cases capture pronounced interpretive divergences.
  • The comparison focuses on codes because they are fundamental to thematic analysis and more specific to utterances than themes or subthemes.
  • Figures 1 and 2 provide visual comparisons of clustering consistency and human–LLM theme assignments.

A. Human and LLM themes

The human and LLM analyses produce different thematic structures for interpreting interactions involving participants with disabilities.

  • Human themes: The human analysis identifies five themes covering interaction fluency, effort, disability needs, power dynamics, and social attributes.
  • LLM themes: The LLM analysis identifies five themes covering interaction modes, psychological safety, embodied presence, social concerns, and practical disability-related value.
  • Clustering consistency: Observed AMI values were 0.231 at the theme level and 0.225 at the subtheme level, both above the random-assignment distribution.
  • Cross-analysis correspondence: Human and LLM theme labels show strongest correspondence for disability needs with social concerns, and for effort requirements with psychological safety and reduced social burden.
  • Cross-analysis correspondence: The human theme power dynamics is distributed across three LLM themes, indicating weaker alignment for a conceptually diffuse category.

C. RQ2: Semantic Consistency

Semantic consistency is strongest for more specific labels and reveals ethical risks when LLM codes alter or flatten participants’ expressed experiences.

  • Semantic consistency: Semantic similarity is highest on average at the code level and decreases at the subtheme and theme levels.The pattern is attributed to increasing abstraction, synthesis, and contextual reasoning demands at higher levels.
  • Semantic consistency: Code-level consistency is highest for sensitivity to disability needs, while subtheme-level consistency is highest for effort requirements.Fluency of interaction and social attributes also show closely comparable subtheme similarity.
  • Qualitative evaluation: The qualitative evaluation reports high bias and marginalisation risk when LLM codes import unsupported concepts, omit explicit ableism, generalise lived anxiety, or replace relational concerns with functional utility.
  • Qualitative evaluation: Two low-grounding cases extract lexically salient elements rather than codes accounting for an utterance’s overall meaning.The evaluation distinguishes interpretive labelling from paraphrasing.

V. DISCUSSION

LLM-assisted thematic analysis shows partial alignment with human analysis: it captures meaningful structure and moderate semantic similarity, but reliability declines with abstraction and interpretive ambiguity. Qualitative disagreements are systematic, including paraphrasing, flattened disability-specific meanings, and risks of marginalisation.

  • RQ1: LLM clusterings are clearly better than random, indicating that they capture meaningful structural patterns rather than assigning labels arbitrarily.Alignment becomes uneven for conceptually diffuse themes such as power dynamics, which spread across multiple LLM categories.
  • RQ1–RQ2: LLM reliability decreases as thematic analysis moves toward abstract and socially situated concepts.Subthemes can perform worse than themes because they require coherent grouping without the flexibility of broader themes.
  • RQ2: Semantic similarity is highest at the code level and decreases at the subtheme and theme levels.Higher-level analysis requires greater abstraction, synthesis, and contextual reasoning than initial coding.
  • RQ3: LLM-generated codes systematically tend toward paraphrase and description rather than interpretive labelling.This matters because coding is described as a theory-building act involving immersion, abstraction, and progressive meaning-making.
  • RQ3: Replacing the participant’s term “ableism” with the broader “discrimination” flattened a disability-specific account of harm into a generic category.The paper interprets this as semantic sanitisation rather than hallucination, raising concerns about representational fidelity and reframing participant perspectives.
  • Ethical significance: Systematic interpretive shifts can impoverish knowledge and create particular misrepresentation risks for vulnerable populations.These risks are especially salient where analysis concerns lived experience and social meaning.

VI. CONCLUSION

The study finds a bounded role for LLM-assisted thematic analysis in HRI involving vulnerable populations. LLMs can support early-stage analysis, while human-in-the-loop review remains valuable for higher-level interpretations involving identity, power, and lived experience.

  • Study scope: The comparative analysis examines human- and LLM-generated thematic analysis for HRI involving vulnerable populations.The study focuses on disability-related HRI research.
  • Implications: LLMs can support early-stage tasks such as coding and clustering, but human-in-the-loop mechanisms remain valuable for higher-level interpretation.The conclusion presents this as a bounded role for LLM thematic analysis.
  • Implications: Human review is particularly critical for topics involving identity, power, and lived experience, where misrepresentations may have ethical implications.The conclusion also proposes testing hybrid human–AI workflows across additional vulnerable populations and contexts.
Loading 2608.21420v1…