Source-linked AI summary
LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding
Robert Chew, John Bollenbacher, Michael Wenger, Jessica Speer, Annice Kim
TL;DR
Deductive coding is reliable but labor-intensive, motivating the question of whether LLMs can reduce its burden without abandoning traditional content-analysis procedures. The paper presents LACA, evaluates GPT-3.5 against human coders across four data sets, and finds comparable agreement in many cases while identifying conditions requiring caution.
Problem
Deductive coding requires substantial manual reading, interpretation, categorization, and reliability work, creating a time burden for researchers handling large text collections.
Method
LACA combines LLM-assisted codebook development, validity and reliability testing against human coders, and final LLM coding when model outputs are adequate.
Results
Across data sets, GPT-3.5 often achieved agreement comparable to human coders, while randomness tests identified codes where performance substantially underperformed.
Takeaways & Limitations
LACA can reduce coding time and support deductive coding, but it is presented as a tool to accelerate labor-intensive later stages rather than replace qualitative researchers.
Takeaways & Limitations
The evaluation forced binary or mutually exclusive outputs even when GPT-3.5 indicated that coding was inapplicable or lacked sufficient information.
Abstract
from arXiv · showhide
Deductive coding is a widely used qualitative research method for determining the prevalence of themes across documents. While useful, deductive coding is often burdensome and time consuming since it requires researchers to read, interpret, and reliably categorize a large body of unstructured text documents. Large language models (LLMs), like ChatGPT, are a class of quickly evolving AI tools that can perform a range of natural language processing and reasoning tasks. In this study, we explore the use of LLMs to reduce the time it takes for deductive coding while retaining the flexibility of a traditional content analysis. We outline the proposed approach, called LLM-assisted content analysis (LACA), along with an in-depth case study using GPT-3.5 for LACA on a publicly available deductive coding data set. Additionally, we conduct an empirical benchmark using LACA on 4 publicly available data sets to assess the broader question of how well GPT-3.5 performs across a range of deductive coding tasks. Overall, we find that GPT-3.5 can often perform deductive coding at levels of agreement comparable to human coders. Additionally, we demonstrate that LACA can help refine prompts for deductive coding, identify codes for which an LLM is randomly guessing, and help assess when to use LLMs vs. human coders for deductive coding. We conclude with several implications for future practice of deductive coding and related research methods.
1 Introduction
Deductive coding systematically classifies text using predefined categories, but becomes burdensome as researchers read, interpret, code, refine codebooks, and assess reliability across large or nuanced data sets. The paper proposes LACA to incorporate LLMs into this workflow and evaluates GPT-3.5 across diverse deductive coding tasks.
- Deductive coding applies theory-based categories and definitions to classify sampled text, contrasting with inductive coding.
- Coding large or nuanced text collections requires repeated reading, careful interpretation, codebook refinement, coder training, and inter-rater reliability assessment.
- LACA incorporates LLMs into deductive coding while retaining alignment with traditional content analysis and helping assess when to use LLMs versus human coders.
- The study demonstrates LACA through an in-depth case study and a benchmark across four diverse publicly available data sets, revealing performance differences across documents, codebooks, and categories.
- LLMs offer zero-shot and few-shot learning capabilities that address a limitation of traditional supervised learning approaches for deductive coding.
2.1 LLM-Assisted Content Analysis (LACA)
LACA modifies traditional deductive coding by using an LLM to co-develop and validate the codebook, assess reliability against human coders, and conduct final coding when performance is adequate. The workflow also uses hypothesis tests, reasoning review, calibration, and prompt refinement to identify and address underperforming codes.
- LACA differs from traditional deductive coding through codebook co-development, human-LLM reliability testing, and LLM replacement of manual final coding.
- Researchers test an LLM-generated codebook on uncoded documents to assess whether the model understands the task and captures the intended construct.
- Hypothesis tests identify codes whose model decisions resemble random guessing, while model-generated reasons support face-validity review and interpretation of decisions.
- After validity testing, a calibration sample is coded by the model and human coders to calculate human-model reliability and compare it with human-human agreement when available.
- When model codes underperform, researchers can refine prompts and codebooks or use alternative coding approaches for affected categories.
- If model codes are non-inferior to human responses, LACA can use the LLM to code all documents or a large random sample, reducing coding burden.
2.2 Experiment
The experiment evaluates GPT-3.5 for deductive coding across four public data sets, using a detailed Trump Tweets case study and broader benchmark comparisons with human coders.
- Experimental design: GPT-3.5 was compared with human coders across four publicly available data sets using temperature 0 to reduce response variability.The study used gpt-3.5-turbo for the comparisons.
- Prompts: The prompts supplied GPT-3.5 with each codebook and text, using single-code prompts for mutually exclusive schemes and code-by-code prompts for non-mutually exclusive schemes.Most coding tasks were zero-shot beyond examples already present in the codebooks.
- Data: The four data sets covered Trump Tweets, Contrarian Claims, BBC News, and Ukraine Water Problems, with coding schemes differing in categories and exclusivity.The Ukraine Water Problems texts included translations from Ukrainian that sometimes introduced imprecision and ambiguity.
- Case study: The evaluation combined a 100-document Trump Tweets development set with qualitative review, randomness tests, and model-generated coding reasons.The case study examined both compelling reasons and hallucinations, defined as unverifiable statements presented as facts.
- Calibration: The calibration set reused the same 100 observations with original and independently replicated human codes to compare human-human and human-LLM agreement.Gwet’s AC1 was preferred because rare codes can produce misleading values for other agreement metrics.
- Summary benchmark: The summary benchmark sampled 100 documents per data set and assessed randomness, inter-rater agreement, and coding time for humans versus GPT-3.5.Binary tasks used two-sided binomial tests, whereas multiple-category tasks used chi-squared tests against equal-probability multinomial distributions, with α = 0.05.
3 Results
Across the Trump Tweets case study and four-dataset benchmark, GPT-3.5 generally achieved agreement comparable to human coders, while randomness tests and coding reasons exposed important failure modes and calibration opportunities. GPT-3.5 also coded documents faster than humans overall, with differences depending on task complexity and prompt design.
- Tests of Randomness: All randomness tests were significant except HSTG, ATSN, CAPT, and INDV, suggesting possible random guessing for those four Trump Tweets codes.The four exceptions included three syntax or formatting codes and one code for references to individuals.
- Inter-rater Reliability: Gwet’s AC1 for human-model agreement was generally high (>0.76) and comparable to human-human agreement, except for codes that failed randomness tests.For the failed codes, human-human agreement was much higher, including 0.96 for HSTG, 1.00 for ATSN, 0.93 for CAPT, and 0.79 for INDV.
- Reasons for Coding Decisions: Coding reasons revealed both overly liberal and overly strict interpretations of the codebook, identifying concrete opportunities to refine prompts and codebook definitions.Examples included treating a political family as immediate family, excluding a specific law-enforcement branch from police, and excluding DACA from U.S. immigration.
- Reasons for Coding Decisions: Model reasoning also identified possible human-coder misunderstandings, including cases where the model’s interpretation was supported by references to the DNC, Melania Trump, or DACA.These examples included disagreement with one human coder and disagreement with both human coders.
- Summary Benchmark Results: Across four datasets, randomness tests generally passed, but two Ukraine Water Problems codes failed despite apparently valid reasoning and low agreement levels.The authors suggest those tests may have failed because coding prevalence was near 50%, rather than because the model was randomly guessing.
- Summary Benchmark Results: Human-model agreement exceeded human-human agreement for three of five Ukraine Water Problems codes and across BBC News codes, but was lower for Contrarian Claims and two Ukraine codes.The benchmark therefore showed task-dependent relationships between model and human agreement.
- Coding Time: 144 sec/doc for humans versus 4 sec/doc for GPT-3.5 was the largest coding-time gap, occurring on the more burdensome Contrarian Claims task.GPT-3.5 also produced coding reasons, while the human coders provided only codes; API requests were made serially.
- Coding Time: 72 sec/doc for humans versus 52 sec/doc for GPT-3.5 was the closest coding-time comparison, partly because Trump Tweets were short and less nuanced.Separate requests for each of 13 codes increased GPT-3.5’s time on this task.
4 Discussion
The study presents LACA as a holistic approach for integrating LLMs into deductive coding, evaluating coding agreement, coding time, randomness tests, and model-generated reasons across multiple datasets. Results suggest GPT-3.5 often reaches human-comparable agreement, while randomness tests and qualitative review help identify disagreements and model limitations.
- LACA integrates LLMs into deductive coding through a case study and evaluation across four publicly available datasets.The evaluation compares coding results and coding time between humans and a commercial LLM.
- Randomness tests can identify codes with substantial human-model disagreement without requiring human-coded data, but manual review remains necessary.The tests may fail when good model performance coincides with a true base rate near equal probability.
- Formalized hypothesis testing extends prior comparisons with random chance into an explicit assessment procedure for LLM coding.The authors describe this as the first work, to their knowledge, to formalize the comparison with hypothesis testing.
- Failed randomness tests revealed formatting-code and other issues linked to GPT-3.5’s byte-pair tokenizer and character-level limitations.The authors state that these issues cannot be fully addressed through additional prompt engineering.
- Model-generated reasons help researchers assess performance, prompt quality, hallucinations, reasoning errors, and codebook definitions.They can also support reflection on human coding decisions and improve instruction readability, although this may vary across models.
- GPT-3.5 often achieves coding agreement comparable to human coders across datasets.Cases of substantial underperformance were usually detectable early through randomness hypothesis tests.
- The authors position LACA as a tool for accelerating manually taxing stages of deductive coding rather than replacing qualitative researchers.They also recommend reporting prompts, model details, codes, reasons, and relevant text to support reproducibility and critique.
- Randomness tests and model-generated reasons can support LLM-labeled data generation without requiring prior labeled data.
Funding
The work was solely funded by RTI International.
- The study was solely funded by RTI International.
Authors’ Contributions
The authors divided study design, funding, prompt development, analyses, data management, qualitative coding, coordination, and writing across the research team.
- RC and AK designed the study and secured funding.
- JB, MW, and RC developed prompts, ran analyses, and managed data.
- JS coordinated qualitative coding and coded documents.
- RC led the writing, with contributions from JB, MW, JS, and AK.
- All authors approved the final manuscript.
A Prompts
The study used dataset-specific prompts for Trump Tweets, BBC News, Ukraine Water Problems, and Contrarian Claims. Each prompt includes codebook, document-text, and output-category components.
- The prompts use codebook and document-text blocks, with initial references that establish the expected input format and available output categories.
- Figure 2 presents the prompt used for the Trump Tweets dataset.
- Figure 3 presents the prompt used for the BBC News dataset.
- Figure 4 presents the prompt used for the Ukraine Water Problems dataset.
- Figure 5 presents the prompt used for the Contrarian Claims dataset.
B Codebooks
The Trump Tweets Codebook operationalizes tweet-level indicators spanning format, references, political content, and climate-related claims. It specifies inclusion and exclusion rules for applying these codes consistently.
- Tweet format: The codebook identifies tweet-format features such as hashtags, @ mentions, capitalized words, and hyperlinks.Links are excluded from hashtag, @-mention, and capitalization decisions; acronyms alone do not count as capitalized-word emphasis.
- Political content: Political-content codes capture criticism, partisan or ideological labels, immigration, and Trump’s campaign slogan.The rules distinguish explicit terms and examples from broader or indirect language, such as “the other side” or general references to America and greatness.
- References: Reference codes cover individuals, immediate family, police, international topics, marginalized groups, and news media.Individual references exclude Trump’s self-references, while marginalized-group coding generally requires a group reference rather than only one individual.
- Climate claims: Climate-related codes classify claims about ice, cooling, warming, sea-level rise, extreme weather, naming changes, natural variation, and climate forcings.Additional categories address greenhouse effects, carbon dioxide, atmospheric CO2, ocean pH, and climate sensitivity.
- Climate claims: The codebook includes categories for non-greenhouse-gas forcings, absent greenhouse-effect evidence, CO2 trends, emissions, and low climate sensitivity.These categories encode specific contrarian claims about climate mechanisms and observations.
C Supplemental Results
The supplemental-results section contains a table presenting detailed inter-rater reliability results.
- Supplemental results: Table 13 presents detailed inter-rater reliability results.