Source-linked AI summary
ChatGPT-4 Outperforms Experts and Crowd Workers in Annotating Political Twitter Messages with Zero-Shot Learning
Petter Törnberg
TL;DR
This paper evaluates ChatGPT-4’s zero-shot classification of U.S. politicians’ Twitter affiliations against expert and crowd-worker annotation. It finds higher accuracy and reliability, with equal or lower bias, and discusses implications and limitations for large-scale interpretive research.
Problem
The paper examines whether ChatGPT-4 can accurately, reliably, and without greater bias classify political affiliation from tweet content compared with human annotators.
Method
The study uses filtered tweets from U.S. senators during the two months before the 2020 election, whose known affiliations provide ground truth for comparison with experts and crowd workers.
Results
ChatGPT-4 achieves higher accuracy and reliability than human classifiers, while its bias is equal to or lower than that of the human groups.
Takeaways & Limitations
LLMs may enable large-scale, replicable interpretive research with lower costs and barriers to entry than conventional text-analysis methods.
Takeaways & Limitations
LLM zero-shot capabilities and their scope remain poorly understood, while performance is sensitive to prompt design and models may not communicate uncertainty.
Abstract
from arXiv · showhide
This paper assesses the accuracy, reliability and bias of the Large Language Model (LLM) ChatGPT-4 on the text analysis task of classifying the political affiliation of a Twitter poster based on the content of a tweet. The LLM is compared to manual annotation by both expert classifiers and crowd workers, generally considered the gold standard for such tasks. We use Twitter messages from United States politicians during the 2020 election, providing a ground truth against which to measure accuracy. The paper finds that ChatGPT-4 has achieves higher accuracy, higher reliability, and equal or lower bias than the human classifiers. The LLM is able to correctly annotate messages that require reasoning on the basis of contextual knowledge, and inferences around the author's intentions - traditionally seen as uniquely human abilities. These findings suggest that LLM will have substantial impact on the use of textual data in the social sciences, by enabling interpretive research at a scale.
ChatGPT and the rise of Generative AI
ChatGPT is a generative AI system whose emergent contextual reasoning enables zero- or few-shot interpretation of text. The paper situates this capability alongside important limitations and growing social-science applications.
- ChatGPT and the rise of Generative AI: ChatGPT is an OpenAI AI chatbot based on GPT-3.5 and GPT-4 language models, trained on extensive text and fine-tuned with supervised and reinforcement learning.Human trainers supplied simulated conversations, and annotators ranked generated responses during reinforcement learning.
- ChatGPT and the rise of Generative AI: Large language models developed emergent capabilities beyond autocomplete, including contextual reasoning that supports zero- or few-shot interpretation.The paper identifies contextual reasoning as the most relevant emergent capacity for its text-analysis task.
- ChatGPT and the rise of Generative AI: ChatGPT remains prone to hallucination, weak symbolic-logic performance, and reproducing biases and racism from its training data.The paper notes that guardrails were imposed to reduce offensive outputs.
- ChatGPT and the rise of Generative AI: Recent social-science studies have applied LLMs to political text tasks including sentiment analysis, ideological scaling, topic modeling, and political argument.
Method: Using ChatGPT to classify Twitter messages
The study evaluates ChatGPT-4 on politically classifying tweets from U.S. senators using known affiliations as ground truth, and compares it with experts and quality-controlled MTurk workers. It uses repeated API runs and assesses accuracy, reliability, and bias against manual annotation.
- Method: Using ChatGPT to classify Twitter messages: The dataset contains 2020-election tweets from U.S. senators, filtered to retain messages at least 100 characters and remove retweets, replies, and URLs.Tweets were collected from September 3 to November 3, 2020, and politicians’ known affiliations provide ground truth.
- Method: Using ChatGPT to classify Twitter messages: ChatGPT-4 was prompted to guess whether each politician was a Democrat or Republican using knowledge of U.S. politics.The instruction required one of the two party labels and told the model to make its best guess when information was insufficient.
- Method: Using ChatGPT to classify Twitter messages: 5,000 model runs varied temperature between 0.2 and 1.0, using five repetitions at each setting to capture stochastic response variability.
- Method: Using ChatGPT to classify Twitter messages: Each question received answers from 10 independent MTurk workers, with U.S.-based Master Qualified workers, control questions, and statistical accuracy checks used for quality control.The study notes that poorly controlled MTurk accuracy was not significantly different from random in pre-study testing.
- Method: Using ChatGPT to classify Twitter messages: Two political-science researchers manually classified all 500 messages to provide an expert comparison, while MTurk plurality responses represented the crowd aggregate.The comparison targets accuracy, reliability, and bias across ChatGPT, crowd workers, and experts.
Results
ChatGPT-4 outperformed human classifiers in accuracy and reliability while showing bias comparable to experts and lower than MTurk workers. It also handled tweets requiring contextual political knowledge and inference about authorial intent.
- The LLM outperformed all individual human classifiers and the combined MTurk crowd in classification accuracy.The comparison includes model runs, individual workers, experts, and the plurality response of 10 MTurk workers.
- The LLM showed much higher intercoder reliability than human classifiers, especially at lower temperature settings.Reliability was measured with Krippendorf’s Alpha; Figure 2 reports 95% bootstrap confidence intervals.
- All classifiers were biased toward guessing Democrat, but LLM and expert bias did not differ significantly, whereas MTurk classifiers were significantly more biased.The comparison examines the fraction of responses assigned to each party.
- ChatGPT correctly classified an implicit tweet about Amy Coney Barrett’s nomination as Republican, as did 7 of 10 MTurk workers.The reasoning linked the timing and wording to Barrett’s nomination and its support by Republicans.
- ChatGPT and 9 of 10 MTurk workers correctly classified a Bible quote as Republican by inferring the political significance of its religious framing.The model’s explanation connected public emphasis on religious values with Republican voters while acknowledging that both parties can be religious.
- The model’s reasoning was plausible and convincing at face value, although the processes underlying its responses remain unknown.The paper separates the apparent quality of the reasoning from uncertainty about how the model generated it.
Discussion
ChatGPT-4 performs well on political-affiliation classification, with findings suggesting higher accuracy and reliability and equal or lower bias than human classifiers. The paper highlights large-scale interpretive research as a potential consequence while emphasizing validation and unresolved limitations.
- Findings: ChatGPT-4 may outperform crowd-workers and expert classifiers in accuracy, reliability, and bias, including tasks requiring contextual and intentional reasoning.The authors report higher accuracy and reliability, with lower or equal bias than standard human approaches.
- Implications: LLMs could enable large-scale, replicable interpretive research with lower costs and barriers to entry.The paper also describes LLMs as less costly and time-consuming than conventional text-analysis methods.
- Caveats: Zero-shot and few-shot capabilities are emergent, poorly understood, and sensitive to prompt design and task formulation.The authors recommend iterative experimentation and validation because LLMs may not disclose uncertainty when they struggle.
- Implications: LLMs may transform social-scientific research while raising epistemological challenges because their capacities are emergent and black-boxed.The paper calls for technical exploration, rigorous research standards, and engagement with these epistemological implications.
4 | Törnberg
The supplied passages contain bibliographic material rather than substantive discussion for this section.
- Related work: The supplied material lists references concerning emergent abilities, zero-shot reasoning, and ideological estimation with large language models.These references are presented as citations rather than findings discussed in the section.
- Publication information: The supplied material also includes working-paper publication metadata.No substantive claim about the paper's argument is provided in this passage.