Source-linked AI summary
One Small Step for Generative AI, One Giant Leap for AGI: A Complete Survey on ChatGPT in AIGC Era
Chaoning Zhang, Chenshuang Zhang, Chenghao Li, Yu Qiao, Sheng Zheng, Sumit Kumar Dam, Mengchun Zhang, Jung Uk Kim, Seong Tae Kim, Jinwoo Choi, Gyeong-Moon Park, Sung-Ho Bae, Lik-Hang Lee, Pan Hui, In So Kweon, Choong Seon Hong
TL;DR
ChatGPT’s rapid adoption and the growing literature around it create a need for a comprehensive review. This paper surveys its technology, applications, and challenges, then considers its possible evolution toward general-purpose AIGC and AGI.
Problem
ChatGPT’s rapid adoption and the expanding literature about it create a need for a comprehensive survey.
Method
The paper synthesizes ChatGPT’s underlying technology, applications, challenges, and development path from Transformer architecture and autoregressive pretraining to GPT models.
Results
More than 100 million monthly active users adopted ChatGPT within two months of its November 2022 release.
Takeaways & Limitations
The survey provides an outlook on ChatGPT’s possible evolution toward general-purpose AIGC for realizing AGI.
Takeaways & Limitations
ChatGPT sometimes generates wrong or meaningless answers that appear reasonable despite its powerful capabilities.
Abstract
from arXiv · showhide
OpenAI has recently released GPT-4 (a.k.a. ChatGPT plus), which is demonstrated to be one small step for generative AI (GAI), but one giant leap for artificial general intelligence (AGI). Since its official release in November 2022, ChatGPT has quickly attracted numerous users with extensive media coverage. Such unprecedented attention has also motivated numerous researchers to investigate ChatGPT from various aspects. According to Google scholar, there are more than 500 articles with ChatGPT in their titles or mentioning it in their abstracts. Considering this, a review is urgently needed, and our work fills this gap. Overall, this work is the first to survey ChatGPT with a comprehensive review of its underlying technology, applications, and challenges. Moreover, we present an outlook on how ChatGPT might evolve to realize general-purpose AIGC (a.k.a. AI-generated content), which will be a significant milestone for the development of AGI.
1 INTRODUCTION
ChatGPT’s rapid adoption and extensive attention created a need for a comprehensive survey. This survey reviews its background, technology, applications, challenges, and possible evolution toward general-purpose AIGC.
- ChatGPT attracted more than 100 million monthly active users within two months of its November 2022 release.
- More than 500 articles already included ChatGPT in their titles or abstracts, motivating broader scholarly investigation.
- The survey presents background on OpenAI and discusses ChatGPT’s capabilities before reviewing its underlying technology.
- Its technology review covers Transformer architecture, autoregressive pretraining, and GPT’s development from version 1 to version 4.
- The survey highlights applications and challenges including technical limitations, misuse, ethics, and regulation.
- It concludes with an outlook on ChatGPT’s possible evolution toward general-purpose AIGC for realizing AGI.
2 OVERVIEW OF CHATGPT
The overview situates ChatGPT within OpenAI’s research history and product development. It describes ChatGPT’s conversational capabilities, broad text-processing uses, and relationship to competing chatbots and search engines.
- 2.1 OpenAI: OpenAI began as a nonprofit research organization focused on deep learning, reinforcement learning, natural language processing, and robotics.
- 2.1 OpenAI: OpenAI was reorganized as a for-profit company in 2019 while continuing to develop ethical and secure AI alongside commercial applications.
- 2.1 OpenAI: OpenAI developed language models, reinforcement-learning algorithms, software tools, and high-performance computing systems for AI research.
- 2.2 ChatGPT: ChatGPT produces detailed, human-like responses through interactive two-way conversations rather than simply directing users to information.
- 2.2 ChatGPT: ChatGPT handles text summarization, completion, classification, sentiment analysis, paraphrasing, and translation.
- 2.2 ChatGPT: ChatGPT is presented as a competitor to search engines, with Microsoft integrating it into Bing to provide more creative responses.
- 2.2 ChatGPT: The overview compares ChatGPT with LaMDA and BlenderBot, noting differences in bias, output constraints, and conversational engagement.
3 TECHNOLOGY BEHIND CHATGPT
ChatGPT builds on Transformer self-attention and autoregressive pretraining, then extends GPT through increasingly general task capabilities and human-feedback training. Across GPT generations, the technology shifts from task-specific fine-tuning toward task-agnostic use, while retaining important limitations.
- Transformer architecture: Transformer self-attention assigns weights to words to capture dependencies and contextual relationships within an input sequence.Queries, keys, and values are linearly derived from the input; similarity-based weights are normalized and applied to values to produce contextual representations.
- Transformer architecture: Attention computes normalized query–key similarities and aggregates the corresponding value representations into a new contextual representation.The similarity is commonly computed with a dot product, normalized by softmax, and used to weight the values.
- Autoregressive pretraining: Autoregressive modeling factorizes a sequence distribution into conditional distributions, predicting subsequent words from preceding sequence elements.Unlike RNNs, autoregressive models use previous time steps as inputs rather than an RNN hidden state.
- GPT model design: GPT models use decoder-only Transformer architectures with self-supervised learning and autoregressive rather than masked pretraining.GPT and BERT both learn from unlabeled text and can be fine-tuned, but GPT predicts tokens autoregressively while BERT predicts masked tokens using bidirectional context.
- GPT model evolution: GPT-1 outperformed task-specific models on 9 of 12 tasks, including natural language inference, question answering, semantic similarity, and text classification.Its performance on zero-shot tasks demonstrated a high level of generalization.
- GPT model evolution: GPT-2 achieved state-of-the-art results on 7 of 8 tested language-modeling datasets but performed poorly on question answering.The evaluated tasks included commonsense reasoning, reading comprehension, summarization, and translation.
- GPT model evolution: GPT-3 performed many tasks without fine-tuning, gradients, or parameter updates, making it task-agnostic compared with fine-tuning-dependent language models.The passage gives language translation as an example of such a task.
- Human-feedback training and later models: Human feedback increased GPT-3.5 usability, while GPT-4 improved professional-test performance and human-intention following relative to earlier versions.GPT-4 scored in the top 10% on the virtual bar exam, compared with GPT-3.5 in the lowest 10%; GPT-4 responses were preferred on 70.2% of 5,214 questions.
4 APPLICATIONS OF CHATGPT
The survey describes ChatGPT applications across scientific writing, education, and medicine, emphasizing assistance with idea generation, literature review, teaching, assessment, and clinical decision support. Reported studies also show strong but uneven performance, including human-comparable results in some educational and medical tasks alongside limitations in literature-review accuracy and clinical use.
- Scientific writing: ChatGPT supports scientific writing through brainstorming, literature review, data analysis, content generation, and grammar checking.It can stimulate new ideas, expand existing ones, and help researchers focus on core research while delegating less creative work.
- Scientific writing: ChatGPT can assist literature review by finding and explaining relevant research, but one test found only 8 of 50 supplied DOIs correctly published.The survey describes topic-based literature searches and paper explanations, while reporting substantial reliability problems in generated references.
- Education field: In education, ChatGPT is used for personalized tutoring, course-material design, adaptive learning, assessment, homework, tests, essays, and academic support.Reported uses include interactive teacher-student dialogue, tailored programs and publications, automated grading, and helping students understand theories and concepts.
- Education field: 71 ± 2% matched the current module average of 71 ± 5% for ChatGPT-generated short-form Physics essays assessed with an authorized method.The result is reported as evidence of capacity to write short-form Physics essays at the level of the module average.
- Medical field: Medical applications include explaining concepts, answering inquiries, assisting students and patients, and supporting clinical decision-making.Studies report 71.7% overall accuracy across published clinical cases, 87% diagnostic accuracy in one study, and 56.25% accuracy for the worst-performing clinical-decision metric set.
- Medical field: ChatGPT’s medical performance is promising but requires professional oversight because it cannot replace licensed diagnosis and its responses can remain imperfect.The survey also reports that experienced users were more likely to distinguish ChatGPT responses from human answers, while another study found no significant difference between them.
5 CHALLENGES
ChatGPT’s challenges span technical limitations, misuse risks, and ethical concerns. Its outputs can be illogical, inconsistent, incorrect, difficult to distinguish from human writing, and susceptible to harmful or unfair use.
- 5.1 Technical limitations: ChatGPT can generate plausible-sounding but wrong or meaningless answers because training can favor human-like responses over correctness.The survey states that factual errors have been mitigated in ChatGPT Plus but remain a problem.
- 5.1 Technical limitations: ChatGPT’s logic reasoning remains limited: it lacks rational human thinking, cannot reliably reason, and may fail difficult mathematical or arithmetic tasks.The survey also describes missing spatial, temporal, physical, behavioral, and psychological inference capabilities.
- 5.1 Technical limitations: ChatGPT may produce different outputs for the same prompt and is highly sensitive to prompt wording, making prompt engineering important for improving query efficiency.The survey notes that changing prompts can yield significantly different outputs.
- 5.1 Technical limitations: ChatGPT lacks self-awareness, consciousness, emotions, and subjective experience, despite generating coherent text and understanding or creating humor.The survey also notes that self-awareness lacks a widely accepted definition and reliable testing methods.
- 5.2 Misuse cases: ChatGPT can be misused for plagiarism, concealed academic submissions, false information, toxic content, malicious software, and harmful online activity.The survey links these uses to risks for education, public information security, individuals, and society.
- 5.2 Misuse cases: Overreliance on ChatGPT may weaken critical and independent thinking, while its human-trained and human-feedback-adjusted outputs can contain political and other biases.The survey also raises concerns about accountability, transparency, affordability, and the difficulty of identifying AI-generated content.
6 OUTLOOK: TOWARDS AGI
The survey outlines two roadmaps for moving ChatGPT toward general-purpose AIGC and AGI: combining it with specialized tools or developing an all-in-one model. It also discusses technical, societal, and governance implications of this evolution.
- Technology aspect: ChatGPT is primarily strong at text-to-text tasks, while GPT-4 remains limited in handling input modalities beyond images and generating outputs beyond text.These constraints keep it from functioning as a general-purpose AIGC tool.
- Road-map 1: combining ChatGPT with other AIGC tools: Combining ChatGPT with specialized AIGC tools could extend its instruction understanding to tasks such as text-to-image generation.The survey contrasts ChatGPT’s instruction capabilities with tools focused on mapping descriptions to images.
- Road-map 2: All-in-one strategy: An all-in-one strategy would handle multiple AIGC tasks within ChatGPT rather than depending on downstream tools, but it presents difficult training and inference-speed challenges.The proposed use cases include generating music from prompts and optional image inputs.
- Technology aspect: The survey identifies a gradual path in which tool combination may be more applicable initially, while ChatGPT increasingly masters AIGC tasks and reduces external dependence.This path is presented as intermediate between the two roadmaps.
7 CONCLUSION
The conclusion presents the paper as a comprehensive survey of ChatGPT’s technology, applications, challenges, and possible evolution toward GAI and AGI. It aims to provide readers with a concise understanding while encouraging further discussion of AGI.
- The paper surveys ChatGPT’s underlying technology, spanning transformer architecture, autoregressive pretraining, and the technology path from GPT-1 to GPT-4.
- It reviews ChatGPT applications across fields including scientific writing, educational technology, and medicine.
- It discusses technical limitations, misuse cases, ethical concerns, and regulation policies surrounding ChatGPT.
- It offers an outlook on roadmaps for ChatGPT’s evolution toward GAI and on how AGI might affect humankind.
- The survey seeks to provide a quick yet comprehensive understanding of ChatGPT and inspire further discussion on AGI.