Source-linked AI summary
Redefining Qualitative Analysis in the AI Era: Utilizing ChatGPT for Efficient Thematic Analysis
He Zhang, Chuhao Wu, Jingyi Xie, Yao Lyu, Jie Cai, John M. Carroll
TL;DR
Thematic analysis is valuable but time-consuming, while prompt-engineering research has paid limited attention to transparency and complex, open-ended qualitative analysis. This study examines ChatGPT for thematic analysis and develops a prompt-design framework, finding that well-designed prompts, greater transparency, and stronger understanding of LLM capabilities can improve interaction with ChatGPT. It also highlights ethical risks and the need to revisit these conclusions as tools and research contexts evolve.
Problem
Thematic analysis requires substantial manual effort, and existing prompt-engineering research gives limited attention to transparency and complex, open-ended qualitative analysis.
Method
The study examines ChatGPT for qualitative thematic analysis and develops a four-part framework covering task background, task description, processing guidance, and expected output content.
Results
Well-designed prompts, enhanced transparency, prompt guidance, and stronger understanding of LLM capabilities can improve users’ interaction with ChatGPT for qualitative analysis.
Takeaways & Limitations
ChatGPT may support qualitative analysis, but researchers should attend to transparency, prompt design, users’ understanding of LLM capabilities, and potential ethical risks.
Takeaways & Limitations
The study relies predominantly on short-term experiments lasting just over an hour, so its insights and framework may require revision as ChatGPT and related tools evolve.
Abstract
from arXiv · showhide
AI tools, particularly large-scale language model (LLM) based applications such as ChatGPT, have the potential to simplify qualitative research. Through semi-structured interviews with seventeen participants, we identified challenges and concerns in integrating ChatGPT into the qualitative analysis process. Collaborating with thirteen qualitative researchers, we developed a framework for designing prompts to enhance the effectiveness of ChatGPT in thematic analysis. Our findings indicate that improving transparency, providing guidance on prompts, and strengthening users' understanding of LLMs' capabilities significantly enhance the users' ability to interact with ChatGPT. We also discovered and revealed the reasons behind researchers' shift in attitude towards ChatGPT from negative to positive. This research not only highlights the importance of well-designed prompts in LLM applications but also offers reflections for qualitative researchers on the perception of AI's role. Finally, we emphasize the potential ethical risks and the impact of constructing AI ethical expectations by researchers, particularly those who are novices, on future research and AI development.
1 INTRODUCTION
Thematic analysis is widely used but can require substantial manual effort as qualitative datasets grow. This study examines whether ChatGPT’s performance in qualitative analysis can be improved through prompt design and develops guidance from participant experiences.
- Thematic analysis identifies and interprets patterns of meaning but becomes time-consuming with large and complex qualitative datasets.
- The study examines ChatGPT as an instrument for thematic analysis, focusing on its practical advantages and limitations.
- The research asks whether prompt design can enhance ChatGPT’s performance in qualitative analysis tasks and, if so, how.
- Researchers recruited 17 participants, analyzed semi-structured interviews, and collaborated with 13 qualitative researchers to identify challenges and techniques for improving ChatGPT’s efficacy.
- The study developed cueing frameworks and discusses ChatGPT’s scope, human-AI roles, ethical implications, and relevance to qualitative analysis.
2 RELATED WORK
Prior work shows ChatGPT’s versatility and the potential of prompt engineering, but complex qualitative analysis remains underexplored. The study therefore frames prompt design as a human-centered, domain-specific problem involving transparency, trust, and support for junior researchers.
- ChatGPT supports diverse language tasks but can also produce nonsensical or incorrect outputs, underscoring the need to acknowledge its limitations.
- Prompt engineering can improve LLM outputs, yet effective use varies by domain and requires application-specific knowledge.
- Existing prompt-engineering research pays less attention to transparency and explainability in complex, open-ended qualitative analysis.
- Thematic analysis is resource-intensive and interpretive, creating challenges involving subjectivity, replicability, generalizability, and researcher expertise.
- The study proposes a human-centric prompt framework to help junior qualitative researchers leverage ChatGPT while addressing these challenges.
- Transparency, explanations, and domain-sensitive communication are important for trust and effective human-AI collaboration.
3 METHODS
The study used semi-structured interviews and a qualitative coding experiment to examine ChatGPT use, then iteratively refined prompts with participants. The resulting design process informed a framework for qualitative analysis tasks.
- Researchers first conducted a pilot interview study with four participants experienced in qualitative methods and ChatGPT.
- Semi-structured interviews explored participants’ qualitative-analysis experiences, ChatGPT challenges, uses, and strategies.
- The researchers distilled the design solutions into a framework of prompts for qualitative analysis tasks.
- The formal study combined interviews with an experiment in which participants used ChatGPT for qualitative analysis.
- Participants designed prompts independently, after which researchers and participants discussed outcomes and collaboratively refined the prompts.
- The research team used reflexive thematic analysis with a six-step procedure and collaborative coding checks.
4 USERS’ EXPERIENCES AND CHALLENGES WITH CHATGPT
Participants identified transparency, performance, prompt-design difficulty, and review cost as major challenges in using ChatGPT for qualitative analysis. Their experiences also showed that understanding ChatGPT’s capabilities and receiving task-specific guidance could improve use and shift attitudes.
- Participants’ main concerns involved ChatGPT’s transparency, consistency, accuracy, prompt-design difficulty, and the cost of reviewing results.
- Initial skepticism largely reflected uncertainty about how ChatGPT generated its outputs.
- Prompt design required substantial effort because online guidance was excessive, attempts could produce divergent outcomes, and tuning was time-consuming.
- Participants’ limited knowledge of ChatGPT’s capabilities reduced performance and discouraged effective use.
- Participants requested customized, task-specific prompts rather than generic guidance for qualitative research.
- Participants favored a standardized yet flexible framework specifying input formats and expected outputs without excessive detail.
5 ANALYSES OF THE DESIGN PROCESS
The section frames prompt design as central to shaping ChatGPT’s performance in qualitative analysis and summarizes strategies developed from participant and researcher input.
- Prompts strongly influence the quality, coherence, and applicability of ChatGPT’s responses in qualitative analysis.The section treats prompt construction as both a mechanism for expressing researcher intentions and a challenge in LLM use.
- Structured prompts should include core information and strategies aligned with participants’ desired qualitative-analysis outcomes.The summarized strategies cover independent coding, prompt-design methods, and testing approaches.
- The section consolidates participant prompt designs and researcher-proposed strategies in Tables 2 and 3.These tables summarize the design-process outputs rather than reproducing identical prompts.
- The summarized strategies are interconnected and are explained in greater detail in subsequent subsections.
5.1 Explanation of Prompts Provided
The framework improves ChatGPT-assisted thematic analysis by specifying task context, analytical procedures, data formats, output requirements, and prioritization while treating role-play as insufficient by itself. It also addresses transparency by requesting explanations and traceable sources.
- 5.1 Explanation of Prompts Provided: Descriptive task background gives ChatGPT the purpose, expected outcomes, and nuances needed for more targeted responses.
- 5.1.2 Prompts Provided: Focus on Methodology (Goal of Task): More specific task descriptions improve performance, and task definitions should be based on the research question and intended analytical method.Examples include analyzing remote-work data for patterns and themes and using the Job Demands-Resources Model.
- 5.1.3 Prompts Provided: Focus on Analytical Process: Pre-cleaning and formatting data are useful because ChatGPT performs better on formatted inputs, although preparation can be less stringent than in traditional analysis.
- 5.1.4 Prompts Provided: Define the Format of the Inputs: Prompts should describe input type, conversational structure, data structure, length, roles, and complexity to reduce discontinuities and corpus misunderstandings.
- 5.1.4 Prompts Provided: Define the Format of the Inputs: Standardized output formats improve readability, transparency, consistency, and transferability into tools such as Excel.Participants also requested themes alongside relevant excerpts and summaries of their significance.
- 5.1.6 Prompts Provided: Role-Playing: Role-playing can focus ChatGPT on a task, but detailed prompt descriptions can replace or surpass it because role-play alone does not produce consistently stable results.
- 5.1.7 Prompts Provided: Prioritization: Prioritization requirements, including limiting the number of codes, can improve readability and help pinpoint key codes for novice users.The proposed framework is intended to remind novices of traditional analysis processes while applying ChatGPT’s capabilities.
5.1.8 Prompts Provided: Clarification, Transparency and Traceability.
The framework combines transparency requests, line-by-line analysis, iterative prompt refinement, and critical user judgment to make ChatGPT-assisted qualitative analysis more interpretable and adaptable. Iteration can improve alignment with research goals but may also cause the model to lose focus without repeated emphasis on the task.
- 5.1.8 Prompts Provided: Clarification, Transparency and Traceability: Combining transparency requests with other prompt strategies can improve the readability of ChatGPT’s qualitative-analysis outputs.
- 5.1.8 Prompts Provided: Clarification, Transparency and Traceability: Analyzing each response independently was more effective than overall analysis for supporting subsequent in-depth studies and potentially generating more discoveries.The authors retain context through complementary strategies such as prioritization.
- 5.1.8 Prompts Provided: Clarification, Transparency and Traceability: Positive feedback on incentives reflected expectations of consistent outputs, but incentives alone did not noticeably improve quality for one participant.Specific strategies such as standardized formats and original information more clearly improved readability and trust.
- 5.1.10 Iteration of Prompts: Natural-language interaction lets users refine outputs conversationally while retaining critical and creative judgment instead of following fixed commands.
- 5.1.10 Iteration of Prompts: The framework treats repeated prompt refinement and critical evaluation of outputs as essential for aligning LLM results with research objectives.
- 5.1.10 Iteration of Prompts: Multiple iterations may cause ChatGPT to lose focus on the original task or context, so objectives and framework components should be re-emphasized.
- 5.1.11 Robustness: Similar outcomes can arise from different prompt strategies, suggesting that feature relevance and ChatGPT’s robustness affect results more than any single technique.A framework may therefore be more useful than a fixed command because it reduces learning costs while providing a minimum satisfactory performance level.
5.2 Notes on ChatGPT with Different Versions
The study used GPT-3.5 rather than GPT-4.0 primarily for accessibility and reported little difference in task outputs between the versions.
- The study selected GPT-3.5 instead of GPT-4.0 because GPT-4.0 had access limits and required a subscription during the formal study.The authors’ tests also found little difference in outputs for the study’s tasks.
6 USER’S ATTITUDE ON CHATGPT’S QUALITATIVE ANALYSIS ASSISTANCE: FROM NO TO YES
Participants shifted from skepticism to acceptance of ChatGPT for qualitative analysis after a structured prompt framework improved transparency, credibility, usability, and confidence while preserving the need for human validation.
- Participants initially worried about ChatGPT’s transparency, performance, prompt difficulty, and the cost of reviewing outputs.
- Step-by-step prompt design improved output format and content, including data-source attribution and table integration for batch coding.
- The framework expanded participants’ understanding of ChatGPT’s qualitative-analysis applications and strengthened acceptance of the tool.
- The framework increased confidence by making ChatGPT’s outputs more interpretable, verifiable, and connected to specific data sources.
- Participants reported that ChatGPT could support preliminary screening and categorization while allowing them to trace claims back to their sources.
- Despite increased trust, participants maintained that human judgment and double-checking were necessary before presenting research findings.
7 DISCUSSION
The discussion presents a structured prompt framework for making ChatGPT-assisted qualitative analysis more transparent and usable, while emphasizing robustness, human oversight, and ethical risks as adoption increases.
- 7.1 Overcoming the Challenges in Prompt Design: The study identifies prompt-design difficulties, especially for junior researchers, and proposes strategies covering context, methodology, data formats, roles, prioritization, transparency, and expertise.
- 7.1 Overcoming the Challenges in Prompt Design: The framework aims to elicit more interpretable and verifiable responses while improving junior researchers’ understanding of ChatGPT’s qualitative-analysis capabilities.
- 7.1 Overcoming the Challenges in Prompt Design: The framework may improve thematic-analysis efficiency, but human oversight and validation remain important in qualitative analysis.
- 7.2 The Robustness of ChatGPT: ChatGPT showed robustness through natural-language understanding and tolerance of spelling or grammatical errors during user interactions.
- 7.3 ChatGPT as a Learning Tool: The prompt framework supported junior researchers by breaking thematic analysis into manageable steps and encouraging more critical reflection on AI outputs.
- 7.4 Future Applications and Innovations: The framework’s flexibility may support new qualitative-analysis applications and interdisciplinary collaboration as more junior researchers use AI-assisted tools.
- 7.5 The Evolving Landscape of Ethical Considerations: Participants’ movement from skepticism toward acceptance highlighted ethical concerns involving transparency, accountability, training-data bias, and reduced emphasis on human critical reflection.
8 LIMITATIONS AND FUTURE WORK
The study identifies boundaries on its exploratory findings and proposes future work to test ChatGPT across more diverse corpora, longer interactions, broader qualitative tasks, and ethical and culturally responsive applications.
- Corpus Scope and Diversity: The selected corpus may not capture the breadth of qualitative data, limiting the generalizability and applicability of the findings.The authors propose incorporating diverse qualitative corpora in future research.
- Long-term Implications of ChatGPT Interactions: The findings are based predominantly on short-term experiments lasting just over an hour, so longer and iterative ChatGPT interactions remain outside the study’s scope.The authors treat the framework as foundational and note that it may require refinement as ChatGPT and newer tools evolve.
- Performance Across Qualitative Analysis Tasks: The research primarily examines coding, leaving ChatGPT’s effectiveness across the broader spectrum of qualitative-analysis tasks largely untested.The authors call for studies covering these uncharted tasks to develop a more holistic understanding of LLMs in qualitative research.
- Future Work: Future directions include extending LLM applications to the humanities, examining ethical implications, and developing personalized tools sensitive to regional and cultural contexts.The proposed toolkit could combine personalized knowledge bases with qualitative-analysis support, while future model updates may affect task performance.
- Future Work: ChatGPT’s demonstrated proficiency in qualitative analysis may extend to quantitative analysis, programming, and creative writing, but these broader applications require further study.The authors frame these as extended applications and capabilities rather than established findings of this study.
- Future Work: Future work should examine ChatGPT’s roles as a tool or co-researcher, including the risks of AI over-reliance and questions about AI understanding.The authors also stress that recognizing ChatGPT’s strengths and limitations is important for advancing AI-assisted qualitative analysis.
9 CONCLUSION
The study first identifies risks and challenges in using ChatGPT for qualitative analysis, then develops a prompt-design framework with qualitative researchers. It reports that transparency, prompt guidance, and understanding LLM capabilities improve interaction and can reverse negative attitudes toward ChatGPT.
- The study used a pilot study to identify risks and challenges associated with ChatGPT in qualitative analysis.
- The researchers examined qualitative analysts’ attitudes through interviews and experiments and collaboratively developed a well-received prompt-design framework.
- Enhancing transparency, providing prompt guidance, and strengthening users’ understanding of LLM capabilities improved interaction with ChatGPT and reversed negative attitudes toward its research use.
- The discussion addresses ChatGPT’s challenges, potential, effects on novice researchers, and ethical considerations in qualitative analysis.
A APPENDIX
The appendix illustrates participant-generated ChatGPT outputs and prompt strategies, including structured tables, Excel transfer, priority requirements, and examples from individual participants.
- P15 specified a table with columns for theme name, frequency, supporting quotes, and the commenting participant.
- One appendix figure demonstrates how a ChatGPT-generated table can be transferred to Excel.
- Another figure presents ChatGPT outputs after participants added priority requirements.
- A further figure shows ChatGPT outputs generated from P8’s prompts.
- The appendix compares outputs obtained by P5 and P6, presenting P5’s result on the left and P6’s on the right.