Source-linked AI summary

Antipatterns in AI-assisted Qualitative Data Analysis: A Catalog of Temptations and Pitfalls for Software Engineering Researchers

Rashina Hoda, Carolyn Seaman, Victoria Gomes, Rodrigo Spinola

arXiv:2608.27927v1cs.SEcs.AI

TL;DR

AI-assisted QDA offers SE researchers greater speed, scale, and reduced manual effort, but researchers lack strategic guidance on its methodological risks. This paper presents a catalog of eleven antipatterns organized into three categories to support responsible analysis and review.

  • Problem

    SE researchers lack strategic guidance for identifying and mitigating methodological risks in AI-assisted QDA despite its expanding opportunities and temptations.

  • Method

    The paper presents a catalog of eleven AI-assisted QDA antipatterns organized into Dangerous Drivers, Operational Missteps, and Analytical Failures.

  • Results

    The catalog provides a structured vocabulary and analytical foundation for reasoning about problematic AI-assisted QDA practices and their risks.

  • Takeaways & Limitations

    The antipatterns help SE researchers identify common temptations and pitfalls, while giving reviewers criteria for recognizing problematic and failed practice.

  • Takeaways & Limitations

    The taxonomy and described relationships are not definitive or exhaustive, and their relevance may change as AI systems evolve.

Abstract

from arXiv · show

AI-assisted qualitative data analysis (QDA) offers unprecedented opportunities to streamline software engineering (SE) research, yet uncritical use risks compromising analytical rigor and flooding the field with accelerated production of low-quality research. While tactical best practices will naturally evolve over time, SE researchers currently lack strategic guidance to identify and mitigate methodological risks when attempting AI-assisted QDA. Based on our decades of qualitative SE research expertise and experience combined with an understanding of the emerging landscape of AI-assisted QDA, this paper presents a catalog of antipatterns in AI-assisted QDA - a set of assumptions and practices that initially appear advantageous but ultimately undermine analytical rigor and validity. The antipatterns are grouped into three categories reflecting escalating impact: Dangerous Drivers, Operational Missteps, and Analytical Failures. As more SE researchers attempt AI-assisted QDA, these antipatterns will help them identify and avoid common temptations and pitfalls, while reviewers can be equipped with the vocabulary and criteria to call out problematic and failed practice. Ultimately, this catalog of antipatterns can serve as a stepping stone in our responsible methodological evolution toward principled and meaningful human-AI collaboration in qualitative research.

1 INTRODUCTION

AI-assisted QDA could expand the speed and scale of SE research, but uncritical use threatens the grounding, reflexivity, iteration, and transparency required for rigorous qualitative analysis. The paper responds by cataloging antipatterns to help researchers and reviewers recognize and mitigate these risks.

  • QDA supports rich, contextualized insights into complex socio-technical phenomena in software engineering research.
  • Core QDA principles include grounding in data, iterative and systematic analysis, researcher reflexivity, and analytical transparency.
  • LLM-based AI promises faster analysis, less manual effort, and processing of larger datasets, creating opportunities and temptations to expand QDA’s scope and scale.
  • Uncritical AI use can obscure data grounding, reduce iterative engagement, diminish reflexivity, and weaken transparency, potentially producing low-quality qualitative research faster and at scale.
  • The field faces simultaneous principled resistance to AI in reflexive qualitative research and rapid promotion of AI-based QDA tools and workflows.
  • Because AI-assisted QDA lacks an established playbook, the paper offers a catalog of eleven antipatterns organized as Dangerous Drivers, Operational Missteps, and Analytical Failures.
  • The catalog provides researchers with warnings and guidance while giving reviewers vocabulary and criteria for identifying problematic practices, although it is neither exhaustive nor final.

2 EMERGING LANDSCAPE OF AI-ASSISTED QDA

The emerging AI-assisted QDA landscape combines longstanding computational support with rapidly expanding LLM applications, but evidence and methodological positions remain varied. Current discussions emphasize preserving human analytical authority, contextual interpretation, reflexivity, and transparency when AI supports selected activities.

  • AI-assisted QDA builds on computational support including CAQDAS, topic modeling, clustering, classification, sentiment analysis, machine learning, and natural language processing.
  • Evidence about these risks is classified as empirically validated, experience-based, unvalidated-assumption, or opinion-based.
  • A group of 419 experienced qualitative researchers from 32 countries rejected generative AI for reflexive qualitative research, grounding their critique in human meaning-making, subjectivity, situated interpretation, and reflexive accountability.
  • Other studies take a cautious, less prohibitive position by examining AI support for selected analytic activities while preserving human analytical authority.
  • SE researchers face practical and methodological tensions because they work with large volumes of socio-technical textual data that still require contextual and interpretive analysis.
  • LLMs have been explored for QDA to reduce manual effort and support analysis, while researchers emphasize that contextual understanding, interpretive judgment, and reflexivity cannot be directly delegated.
  • Guidance for empirical SE research recommends reporting model versions, prompts, parameters, and usage settings because QDA outputs are sensitive to these factors, model updates, sampling decisions, and nondeterministic behavior.
  • The literature describes recurring AI-assisted QDA risks as antipatterns, extending an established SE vocabulary for identifying what not to do.

3 A CATALOG OF ANTIPATTERNS

The paper catalogs AI-assisted QDA antipatterns: seemingly reasonable assumptions and practices that undermine qualitative rigor and validity. Eleven antipatterns span motivations, operational use, and analytical failures, with cascading risks requiring meaningful human-AI collaboration.

  • Antipatterns are desirable-seeming assumptions and practices that ultimately undermine the rigor and validity of AI-assisted QDA.
  • Dangerous Drivers: Using AI because it is available or prioritizing productivity can create cascading pressure toward greater speed and scale, with losses in depth, reflexivity, traceability, and transparency.
  • Operational Missteps: AI categorization, black-box use, excessive trust, and treating AI output as insight can reduce qualitative analysis to coarse categories, overstate findings, and obscure context.
  • Analytical Failures: Fully automated QDA removes human review and course correction, leaving AI decisions unchecked and driving losses of traceability, reflexivity, expertise, depth, and emergence.
  • Compounding and Cascading Antipatterns: Early methodological weaknesses can propagate through the workflow, progressively distancing researchers from data and compromising grounding, reflexivity, transparency, and trustworthiness.

4 RECOMMENDATIONS

The recommendations translate antipatterns into safeguards for researchers and reviewers, emphasizing human leadership, methodological expertise, holistic diagnosis, and transparent evaluation. They aim to support critical, accountable AI-assisted QDA without discouraging AI use outright.

  • Recommendations for Researchers: Researchers should lead QDA themselves and possess enough QDA and AI expertise to question outputs, verify them, and recognize inappropriate uses.The human-in-the-lead approach preserves the analyst’s role in shaping interpretation rather than treating AI as the driver and the human as a reviewer.
  • Recommendations for Researchers: Researchers should examine motivations for adopting AI, especially peer pressure and fear of missing out, before analysis begins.Unexamined motivations can initiate antipatterns that later appear as operational missteps or analytical failures.
  • Recommendations for Researchers: Researchers should address upstream assumptions because antipatterns can cascade and compound across stages of the research process.Preventing early temptations and missteps may reduce the likelihood of more severe downstream failures.
  • Recommendations for Researchers: Researchers should recognize that repeated antipatterns can degrade qualitative research and erode the community’s ability to evaluate AI-generated results.Outsourcing analytical skills to AI may weaken qualitative researchers’ methodological repertoire and their ability to act as effective arbiters.
  • Recommendations for Reviewers: Reviewers should evaluate how AI introduced risks into the analytical process rather than treating AI use itself as the central criterion.The catalog supplies vocabulary for identifying problematic assumptions, missteps, and failures and for giving precise feedback on credibility.
  • Recommendations for Reviewers: Reviewers should flag fully automated pipelines, missing chains of evidence, minimal raw-data engagement, and findings not grounded in data.These patterns indicate Analytical Failures and are described as methodologically unsound qualitative research.
  • Recommendations for Reviewers: Reviewers should assess sustained human engagement, traceability, transparency, and accountability throughout the analytical process.They should examine how AI was used, how outputs were generated, and how outputs were interpreted in relation to the data.
  • Recommendations for Reviewers: Reviewers should use the antipatterns diagnostically and evaluate studies holistically because one manifestation may signal other apparent or hidden antipatterns.Specific labels can make methodological concerns more systematic and actionable than general impressions.

5 FUTURE DIRECTIONS

The paper proposes moving from reactive identification of AI-assisted QDA antipatterns toward intentional use, principled human–AI collaboration, and systematic evaluation. These directions support a broader agenda for responsible QDA evolution while preserving qualitative research’s methodological integrity.

  • Improve and articulate intentions: Future research should clarify when and why AI belongs in QDA and how its use aligns with research objectives.This includes distinguishing superficial from substantive contributions and moving beyond adoption driven by convenience, novelty, or perceived efficiency.
  • Design principles for human-AI collaboration: Meaningful human–AI collaboration should define interaction models, analytical responsibilities, and workflows that preserve interpretation and reflexivity.AI should support rather than replace core analytical processes, avoiding practices such as the black box trap and automated QDA.
  • Develop systematic evaluation guidelines: Future work should develop evaluation and reporting guidelines that explicitly account for AI’s role in the analytical process.Such guidance addresses limited current direction and inconsistent, potentially extreme or superficial evaluation practices.
  • Develop systematic evaluation guidelines: Consistent, transparent, and rigorous evaluation mechanisms are essential for assessing human–AI collaboration in qualitative research.
  • Together, these directions establish a research agenda for purposeful and meaningful AI use that harnesses AI’s potential while preserving qualitative research’s methodological integrity.Identifying antipatterns is presented as a first step toward principles, guidelines, and exemplars of valuable use.
Loading 2608.27927v1…