Source-linked AI summary
Decoding ChatGPT: A Taxonomy of Existing Research, Current Challenges, and Possible Future Directions
Shahab Saquib Sohail, Faiza Farhat, Yassine Himeur, Mohammad Nadeem, Dag Øivind Madsen, Yashbir Singh, Shadi Atalla, Wathiq Mansoor
TL;DR
ChatGPT research has expanded rapidly across applications, but bias, trustworthiness, reliability, and ethical-use concerns remain unresolved. This paper reviews the literature to build a research taxonomy, examine applications and common approaches, and identify challenges and future directions. It concludes that broader deployment should be accompanied by work on training, human feedback, trustworthy AI, and ethical safeguards.
Problem
ChatGPT’s expanding use is accompanied by persistent concerns about hallucination, misinformation, bias, trustworthiness, ethical use, and overreliance.
Method
The paper conducts a comprehensive review of over 100 Scopus-indexed ChatGPT publications, classifying applications, research approaches, issues, and future directions.
Results
The review identifies diverse applications and classifies ChatGPT limitations into intrinsic and usage-related issues while discussing ethical concerns and future enhancements.
Takeaways & Limitations
Future ChatGPT development should address reliability, fairness, data and usage ethics, human feedback, and trustworthy-AI considerations.
Abstract
from arXiv · showhide
Chat Generative Pre-trained Transformer (ChatGPT) has gained significant interest and attention since its launch in November 2022. It has shown impressive performance in various domains, including passing exams and creative writing. However, challenges and concerns related to biases and trust persist. In this work, we present a comprehensive review of over 100 Scopus-indexed publications on ChatGPT, aiming to provide a taxonomy of ChatGPT research and explore its applications. We critically analyze the existing literature, identifying common approaches employed in the studies. Additionally, we investigate diverse application areas where ChatGPT has found utility, such as healthcare, marketing and financial services, software engineering, academic and scientific writing, research and education, environmental science, and natural language processing. Through examining these applications, we gain valuable insights into the potential of ChatGPT in addressing real-world challenges. We also discuss crucial issues related to ChatGPT, including biases and trustworthiness, emphasizing the need for further research and development in these areas. Furthermore, we identify potential future directions for ChatGPT research, proposing solutions to current challenges and speculating on expected advancements. By fully leveraging the capabilities of ChatGPT, we can unlock its potential across various domains, leading to advancements in conversational AI and transformative impacts in society.
I. INTRODUCTION
ChatGPT is presented as a conversational language model whose rapid adoption has prompted research across disciplines, applications, risks, and future capabilities. This review organizes that literature around ChatGPT’s architecture, diverse uses, unresolved concerns, and possible enhancements.
- Background: ChatGPT uses GPT-based language processing to generate conversational responses and has been applied in customer service, virtual assistants, and related settings.The model processes user inputs and generates replies, with GPT models supporting tasks such as translation, summarization, and question answering.
- Architecture: Reinforcement learning from human feedback helps ChatGPT learn human preferences through extended dialogues, while transformer components process and generate text.The described architecture includes encoder and decoder components, with attention supporting selective processing of input information.
- Applications and concerns: ChatGPT research spans healthcare, cybersecurity, environmental studies, scientific writing, education, and other domains despite concerns about factual accuracy and ethical use.The review identifies applications across multiple fields while noting risks from unreliable, misleading, or copyright-infringing content.
- Research questions: The review asks how ChatGPT research is progressing, how publication trends and applications vary, how multimodal data may enhance it, and how deployment risks can be addressed.Its research questions cover architecture, publication diversity, applications, multimodal capabilities, fairness, transparency, explainability, and human-centered design.
- Contributions: The study presents a ChatGPT-specific taxonomy, surveys eight application areas, examines current issues, and discusses future enhancements and applications.It positions itself as a comprehensive critical study addressing limitations, ethical concerns, and future directions beyond broader LLM and AIGC reviews.
II. SURVEY METHODOLOGY
The survey follows a Kitchenham-based literature-review methodology, applying explicit inclusion and exclusion criteria to Scopus-indexed ChatGPT publications. It ultimately analyzes 109 articles and characterizes their subject areas and international contributions.
- Review design: The review methodology follows Kitchenham (2004) and was motivated by the rapid growth and diversity of ChatGPT research.The study sought to organize applications, limitations, and future directions across the expanding literature.
- Selection criteria: Eligible papers were peer-reviewed, English-language articles discussing ChatGPT, containing the term in the title or abstract, and published by March 25, 2023.The criteria also required the paper’s foundation to be a peer-reviewed publication and excluded works focused only on GPT or generative AI generally.
- Corpus: Medicine accounted for 23% of publications, followed by social sciences at 20% and computer science at 11%.These subject-area proportions were used to characterize the disciplinary distribution of the corpus.
- International participation: The United States produced 33 publications and collaborated with 24 countries, while the United Kingdom produced 10 publications.Australia and China each produced 9 publications; Switzerland, Australia, and the United Kingdom followed the United States in collaboration breadth.
III. DIVERSITY OF PUBLICATION ON CHATGPT
ChatGPT publications rapidly diversified across disciplines, publication venues, research purposes, and countries. The literature was dominated by evaluations, predictions, and reviews, with scientific writing and education among prominent topics.
- Publication landscape: The literature includes applications and discussions spanning content creation, translation, essays, coding, research assistance, and scientific writing.Nature published 13 articles, while Accountability in Research published 4 and several medical-education and digital-health journals published 3 each.
- Research categories: The reviewed corpus mainly comprised 68 ChatGPT evaluations, followed by prediction studies and reviews published through March 25, 2023.The largest group assessed ChatGPT’s accuracy or knowledge depth across domains.
- Common approach: Most studies selected a topic, queried ChatGPT, and interpreted its responses as either future predictions or assessments of ChatGPT’s features and impacts.This recurring structure is identified as a prevalent pattern across the literature.
- Research emphases: Scientific writing was evaluated in 30 publications, while education-related evaluation was the second-largest issue and research capability formed another major group.Studies also assessed bias, study support, mental health, public health, and subject expertise through test questions.
- Prediction studies: Education predictions appeared in 12 publications, alongside forecasts about disease diagnosis, research trends, and academic or scientific writing.The review also found that only one of ten prior reviewed articles discussed AI-driven conversational chatbots including ChatGPT, and none was exclusively a ChatGPT systematic review.
- Emerging trends: Prompt engineering emerged as a related research dimension involving explicit instructions and example outputs to elicit desired responses.The review identifies prompt design as an emerging trend in ChatGPT research.
V. APPLICATIONS OF CHATGPT
ChatGPT is applied across healthcare, financial services, education, environmental science, language, and robotics, with reported benefits alongside concerns about reliability, trust, and safe use.
- Marketing and financial services: In financial services, ChatGPT can support back-end operations, data analysis, and personalized offers, but human involvement remains necessary to verify trustworthiness.
- Healthcare: Medical evaluations report both passing performance and serious weaknesses, including inconsistent advice, omitted clinical factors, and dangerous antimicrobial recommendations.
- Healthcare: ChatGPT supports healthcare tasks including medical analysis, decision support, diagnosis, research-question generation, and interpretation of patient or test data.
- Education, language, and robotics: ChatGPT can assist learners with vocabulary and ambiguity while supporting robot troubleshooting, task allocation, communication, and user guidance.
- Environmental science: ChatGPT may aid environmental data analysis, education, policy, climate modeling, resource management, and monitoring for climate-related applications.
C. Software Engineering
ChatGPT has been used across software engineering subprocesses, particularly conversational coding and automated bug fixing, while architecture analysis benefits from human observation.
- ChatGPT assists software development, design, testing, coding, and other software-engineering subprocesses.
- Conversational code input can make coding more user-friendly and intuitive for less-skilled programmers.
- Researchers have used ChatGPT to automate programming-bug fixes and evaluated its performance against standard methods on the QuixBugs benchmark.
- ChatGPT can analyze, synthesize, and evaluate software architecture when human observation is present.
D. Academic and scientific writing
ChatGPT is widely used for academic and scientific writing, but its usefulness depends on human oversight because writing assistance raises accuracy, plagiarism, authorship, and access concerns.
- ChatGPT has been used for essays, applications, emails, medical articles, and research papers, with authors emphasizing proper use and human mentorship.
- ChatGPT can write in a human style and may reproduce authors’ writing styles.
- Academic chatbot use raises ethical concerns involving plagiarism, inaccuracies, unequal accessibility, and disagreement over AI authorship.
E. Research and Education
Research and education studies report useful performance across technical and non-technical tasks, but they also emphasize limitations, misleading outputs, and the need for careful validation.
- ChatGPT has been experimentally used for technical tasks such as engineering and programming and for non-technical tasks such as language and literature.
- Researchers warn that bias, discrimination, privacy, security, misuse, accountability, and transparency constrain research and education applications.
- ChatGPT performs well on structured tasks such as code translation and explaining familiar concepts but struggles with unfamiliar terms and code generation from scratch.
- AI-assisted tools should be used cautiously, properly validated, and combined with other methods to support accurate software-process improvement.
F. Environmental Science
ChatGPT is presented as a promising but still sparsely studied tool for environmental science, supporting research workflows, climate-data analysis, and public communication.
- ChatGPT may streamline environmental research workflows, allowing researchers to focus more on experiment design and new ideas rather than writing quality.
- ChatGPT could increase representation from non-English-speaking countries in environmental research.
- ChatGPT can help analyze climate-change data, predict climate-shift patterns, simplify complex information, and provide policymakers with relevant information.
G. Natural Language Processing
The reviewed literature portrays ChatGPT as capable across several NLP tasks, but constrained by reasoning, language-resource, factuality, bias, ethical, and detection challenges.
- Applications: ChatGPT has been applied to suicide-tendency, hate-speech, and fake-news detection, with larger models performing some NLP tasks without task-specific data adaptation.
- Evaluation: ChatGPT excelled at arithmetic reasoning but faced challenges in sequence tagging.
- Machine translation: GPT models produced highly competitive machine translations for languages with abundant resources, while performance remained more difficult for low-resource languages.
- Applications: ChatGPT applications are expanding across domains, including news-accuracy assessment and fake-news detection.
- Intrinsic limitations: ChatGPT’s intrinsic limitations include hallucination, biased content, lack of real-time information, misinformation, and inexplicability.
- Intrinsic limitations: Hallucinated or misleading outputs can produce counterfactual or meaningless responses that threaten reliability and may be mistaken for legitimate information.
- Intrinsic limitations: Biases and stereotypes raise concerns in applications requiring accurate information and explicit reasoning, including finance, environmental science, and healthcare.
- Usage-related issues: Usage-related concerns include unethical content generation, copyright infringement, and overreliance on ChatGPT.
VII. FUTURE POSSIBILITIES
The review proposes future improvements to ChatGPT’s conversational ability through broader data, fine-tuning, human feedback, emotional understanding, and stronger contextual comprehension, while noting associated risks.
- 1) Increasing the volume and variety of training data: Broader and more diverse training data could improve ChatGPT’s ability to handle varied topics, contexts, language patterns, and speech patterns.Data quality, relevance, and representativeness remain important conditions for improvement.
- 2) Fine-tuning: Fine-tuning ChatGPT on conversational tasks or specific domains could produce more pertinent, useful, and human-like responses.Examples include customer service and personal-assistant applications.
- 3) Incorporating human feedback: Improving comprehension of conversation history and user intent could make ChatGPT’s responses more pertinent and personalized.Context understanding is presented as a target for incorporating feedback and improving conversational performance.
- 3) Incorporating human feedback: Human ratings, reviews, and feedback can be analyzed and incorporated through reinforcement learning or fine-tuning to improve responses iteratively.The process includes collecting feedback, identifying weaknesses, updating the model, and evaluating the revised version.
- 4) Incorporating human emotions: Adding emotional capabilities could make interactions more relatable, but risks include bias, discrimination, manipulation, privacy evasion, and offensive responses.The review emphasizes accountability, transparency, user permission, and involvement from psychology and ethics experts.
5) Style-based technique for higher level text analysis:
The review discusses style-based analysis, personalization, domain adaptation, cultural awareness, and continuous feedback as ways to make ChatGPT’s responses more stylistically aligned and user-specific.
- Style-based technique for higher level text analysis: Style-based techniques and word embeddings could help ChatGPT identify, replicate, and generate text with desired stylistic patterns.The proposed integration combines stylistic analysis with ChatGPT’s conversational framework.
- Personalization: Personalizing ChatGPT with prior interactions could support more intimate conversations while improving user profiling and data protection.The review identifies personalization as a future adaptation based on users’ previous interactions.
- Personalization: Additional information from social media, customer support, and online conversations could improve linguistic understanding and user-specific recommendations.The proposed data sources aim to provide more diverse language and deeper linguistic nuance.
- Domain adaptation: Fine-tuning on domain-specific datasets could produce more precise and tailored answers in areas such as customer service, healthcare, business, and finance.Domain adaptation is presented as a way to increase ChatGPT’s knowledge for particular users and topics.
- Personalization: Cultural training data, personalized prompts, and user feedback could help ChatGPT produce more respectful, individualized, and accurate responses.The proposed feedback sources include surveys, user trials, and analysis of interactions.
- Continuous training and updating: Continuous training with new data could keep ChatGPT’s individualized responses aligned with emerging topics and trends.One proposed approach is regularly adding new data to the existing training set or fine-tuning the model.
C. Multimodal design
Multimodal ChatGPT would combine text, audio, images, video, and human-computer interaction technologies to support more natural and effective communication, alongside fairness, privacy, and accountability requirements.
- C. Multimodal design: Multimodal integration would combine text, audio, and images to create more intuitive, engaging, and effective user experiences.The review presents multimodality as a future direction for more natural human-computer communication.
- Image-based design: Image recognition, image captioning, and image-based search could extend ChatGPT with visual understanding and visual-content retrieval.These technologies are described as components of image-based multimodal design.
- Audio-based design: Speech recognition, audio captioning, and content-based audio retrieval could help ChatGPT process spoken input and improve accessibility for deaf, hard-of-hearing, and nonnative-speaking users.Speech recognition translates aural input into text that software can process.
- Video-based design: Video analysis, captioning, and indexing could enable multimodal ChatGPT to process footage, detect objects, track motion, and provide captions.These capabilities are presented as components of video-based multimodal systems.
- Human-computer interaction: Facial expressions, touch sensitivity, and biometric recognition could support more natural interaction and improved security.Examples include responding to facial emotions, recognizing touch inputs, and using fingerprint or iris authentication.
- Trustworthiness: Trustworthy multimodal AI requires fairness, ethical data practices, privacy protection, transparency, accountability, and continuous monitoring.The review identifies encryption, differential privacy, and bias monitoring as relevant safeguards.
2) Transparency:
Transparency and explainability are presented as foundations for understanding, validating, and trusting ChatGPT’s decisions, but fairness requires continuous monitoring and improvement.
- 2) Transparency: Transparent AI makes decision-making processes and underlying data more accessible, helping users understand how systems are deployed and monitored.The review links transparency with ethical and legal decision-making.
- 2) Transparency: Explainable AI could justify ChatGPT’s choices, while feature-importance scores, visualizations, interpretable models, and human oversight could clarify its outputs.These approaches aim to show influential inputs and how the system arrives at its final output.
- 2) Transparency: ChatGPT’s statistically patterned responses may sometimes be incomplete or incorrect, making explanation and validation especially relevant.The review presents explainability as a way for humans to understand and validate system decisions.
- VIII. CONCLUSION: The review’s taxonomy covers more than 100 Scopus-indexed publications and applications across healthcare, marketing and finance, environment, education and research, and academic writing.It also identifies intrinsic, usage-centric, and ethical issues and discusses future directions.
- VIII. CONCLUSION: Improving conversational capabilities, personalization, multimodal design, and trustworthiness are identified as promising future directions for ChatGPT.The review connects these directions with efforts to address current challenges and improve efficacy.