Source-linked AI summary

LLMs for Survey Text Analysis - A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis

Leonardo Bergmann, Renata Gheorghiu, Ana Gvritishvili, Alex Mican, Chris Stewart, Topias Tolonen-Weckström

arXiv:2608.22417v1cs.AIcs.CLcs.HC

TL;DR

Evidence for LLM performance in inductive content analysis remains limited, motivating a comparison with human coding. The study applies human and LLM-supported inductive analysis to survey data and reports between-entity agreement for coding and themes.

  • Problem

    Evidence on LLM performance in inductive content analysis and comparison with humans remains limited.

  • Method

    The study compares five human inductive analyses with GPT-5.4 using an established LLM-supported procedure on open-text survey data.

  • Results

    ARI was 0.61 for agreement between humans and the LLM on coding and 0.54 for themes.

  • Takeaways & Limitations

    The study reports alignment between human and LLM outputs, particularly for coding, in this analysis.

  • Takeaways & Limitations

    The study’s discussion identifies limitations affecting interpretation, validation, and theory-building.

Abstract

from arXiv · show

Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet evidence on their performance in inductive content analysis remains limited. This study compares human and LLM-based inductive coding of open-ended survey responses from 903 answers across six variables from a European PhD student survey. Five human coders performed inductive content analysis following a standardized coding scheme, while an LLM (GPT-5.4) conducted the same task using an established prompting procedure. Agreement between human and LLM outputs was assessed using the Adjusted Rand Index (ARI). Results showed an alignment between humans and the LLM, with ARI values of 0.61 for coding and 0.54 for theme generation. These values were close to the internal consistency of coding and theme results within humans (ARI = 0.68) and the LLM (ARI = 0.76). Agreement varied widely across variables, with low within-entity consistency consistently linked to low between-entity agreement, underscoring the role of data characteristics and individual performance in reliability. Overall, the findings suggest that LLMs can approximate human coding in this case-specific setting, particularly at the coding level, and may serve as a scalable support tool for inductive qualitative analysis.

Introduction

The introduction presents content analysis as a systematic approach to extracting categories from text and identifies limited evidence for LLM-supported inductive analysis compared with human coding.

  • Content analysis systematically extracts content-related categories from textual data through deductive or inductive approaches.
  • Inductive analysis generates initial semantic codes from textual data and groups them into emergent themes.
  • Emergent themes can condense original material into constructs that support qualitative, quantitative, and computer-supported analysis.
  • Existing studies provide evidence for reliable LLM performance in simple deductive coding, especially with binary categories.
  • Evidence remains rare for inductive LLM approaches and their comparison with humans.
  • The study compares human and LLM coding results using LLM-supported inductive content analysis and cluster correlation.

Methods

The methods use European PhD-student survey responses to compare five human inductive analyses with an LLM-supported analysis based on an established procedure.

  • The European Students’ Union collected survey data from 2,800 PhD students across Europe.
  • Six open-text survey items were evaluated with an LLM-supported inductive content-analysis procedure.
  • A randomly drawn 10% of survey responses served as the study’s data basis.
  • Five human coders independently performed inductive content analysis using the coding instructions and mask in Appendix A.
  • GPT-5.4 was accessed through the OpenAI API and applied the procedure outlined by Bergmann et al. (2025).

Appendix B.

The appendix describes Adjusted Rand Index comparisons of human and LLM coding and themes, accounting for multi-label code assignments.

  • Adjusted Rand Index compares two partitions by assessing whether pairs of points occupy the same or different clusters.
  • ARI equals 0 for random partitioning and is bounded above by 1 for perfect agreement.
  • The analysis compared coding and theme results between humans and the LLM, as well as within each entity.
  • Because responses could receive multiple codes, assignments were transformed into co-assignment matrices before ARI calculation.
  • Co-assignment matrices enabled comparison of human and LLM outputs despite the multi-label coding structure.
  • The quantitative comparisons were supplemented by asking human coders about their impressions of the LLM results and whether they would trust the LLM for the task.

Results

The results comprise 903 responses across six survey variables, with sample sizes ranging from 55 to 266 responses.

  • 903 survey responses were analyzed across six variables.
  • The full original survey question for each variable is available elsewhere in the paper.

Appendix C.

Agreement between humans and the LLM was 0.61 for coding and 0.54 for themes, while within-entity consistency was 0.68 for humans and 0.76 for the LLM. Agreement varied substantially across variables and was associated with internal consistency.

  • 0.61 ARI was the average agreement between human and LLM coding results, while theme agreement was 0.54 ARI.
  • 0.68 ARI characterized human internal consistency, compared with 0.76 for the LLM.These comparisons concern alignment between coding-level and theme-level classifications generated by the same entity.
  • ARI values across variables ranged from 0.31 to 0.89, indicating substantial variation in agreement.
  • Financial Situation had all ARI values at or below 0.71, whereas Anything Else had all values at or above 0.70.
  • Low internal consistency for the LLM coincided with low theme agreement for Financial Situation and Well Being, while low human consistency coincided with moderate agreement for Personal Experience.
  • The authors describe the overall ARI performance as promising, and human coders reported that the LLM’s categorizations reflected their own.

Discussion

This case-specific study found moderate human–LLM agreement in inductive content analysis, especially for coding, with lower alignment for themes. Agreement varied across variables, and low within-entity consistency was consistently associated with low between-entity agreement.

  • Human–LLM agreement: ARI = 0.61 for coding and ARI = 0.54 for themes measured agreement between humans and the LLM.Theme generation involved higher-level abstractions and more degrees of freedom, which may contribute to lower agreement.
  • Human–LLM agreement: ARI = 0.68 for human coding and theme alignment was close to the human–LLM results, suggesting comparable differences.The authors interpret the observed human–LLM agreement as relatively robust and comparable to typical human consistency.
  • Variation across variables: Agreement varied strongly across variables, with some topics showing high alignment and others much lower alignment.The variation indicates that data characteristics and individual performance contribute to reliability.
  • Implications: The study suggests that LLMs may support early, labor-intensive stages of inductive coding, while human researchers remain essential for contextual interpretation, validation, and theory-building.The conclusion is case-specific and does not support replacing human qualitative judgment.
  • Limitations and future research: Future research should examine different models, prompting strategies, and evaluation approaches because this study relied on one model, lacked clear ground truth, and used single human coders per variable.The authors specifically recommend using more human coders per variable.

Author Contributions (CRediT framework)

The supplied passages combine author-contribution records with instructions and examples for inductive content analysis, including code generation and theme grouping.

  • Author Contributions: Leonardo Bergmann is credited with conceptualization, data curation, formal analysis, methodology, project administration, resources, and the original draft.
  • Author Contributions: Investigation is credited to Leonardo Bergmann, Renata Gheorghiu, Ana Gvritishvili, Alex Mican, Chris Stewart, and Topias Tolonen-Weckström.
  • Author Contributions: Review and editing are credited to Renata Gheorghiu, Ana Gvritishvili, Alex Mican, and Chris Stewart.
  • Inductive Content Analysis: The analysis process generates concise codes for individual responses, reuses existing codes where possible, permits multiple codes, and groups related codes into overarching themes.
  • Inductive Content Analysis: An example maps reduced tuition fees, increased scholarships, and low-income student support to the broader theme “Financial Accessibility.”

Appendix B - LLM-Supported Inductive Content Analysis

The appendix describes a two-stage LLM-supported inductive content-analysis workflow: generating categories from responses, then assigning responses and codes to broader categories.

  • Category Generation: The workflow first extracts recurring themes, topics, or intents from survey responses and produces categories with short descriptions.
  • Implementation: GPT-5.4 was accessed through OpenAI’s API, with prompts and survey responses inserted into the user field; a screenshot documents the model settings.
  • Output Format: Outputs are structured as Markdown tables, with category descriptions for generated categories and message-level classifications for assignments.
  • Category Assignment: The response-level classification stage assigns each message to one or multiple predefined categories and explains each assignment.
  • Category Generation: Category generation then identifies overarching categories from primary categories by finding recurring themes, topics, or intents.
  • Grouping Categories: The grouping stage assigns each primary category to the most relevant overarching category and provides a categorization explanation.

Appendix C - Full Survey Questions

The appendix presents the full survey questions for the six analysed variables, including questions related to doctoral education in respondents’ countries.

  • Survey Questions: The appendix displays the original full survey question for each of the six analysed variables.
  • Survey Questions: The appendix includes Table C.1 alongside the survey-question material.
  • Survey Questions: The displayed questions concern doctoral education in the respondents’ countries.

Appendix D - Human Coder Opinion on LLM Coding

Human coders generally viewed the LLM’s categorization as aligned with their own, while identifying risks of overinterpretation, reduced precision, and overly broad themes.

  • Overall Assessment: Reviewers reported that the LLM grouped ideas effectively and often captured the same main idea, sometimes with shorter, more concrete formulations.
  • Concerns: The model tended to create signals from noise and failed to replicate a code concerning the possibility of having another job, requiring vigilance during interpretation.
  • Variable-Specific Feedback: Financial responses were difficult to evaluate because grants, scholarships, salaries, system-level funding, and grants involve nuanced distinctions.
  • Concerns: A recurring concern was that the LLM overinterpreted short or unclear responses, including coding an unclear response without sufficient basis.
  • Concerns: Reviewers also observed reduced coding precision when the LLM treated distinct human cases identically.
  • Concerns: The LLM sometimes broadened themes by collapsing multiple human codes and themes into one less precise category.
  • Overall Assessment: Human feedback described the LLM analysis as generally fine, broadly aligned with human categorization, and suitable for continuing the remaining analysis.
Loading 2608.22417v1…