Source-linked AI summary
GPTEval: A Survey on Assessments of ChatGPT and GPT-4
Rui Mao, Guanyi Chen, Xulang Zhang, Frank Guerin, Erik Cambria
TL;DR
A comprehensive review of ChatGPT and GPT-4 assessments is needed because prior studies span many tasks and disciplines without a collective synthesis. This survey analyzes their language, reasoning, scientific, and ethical assessments, compares findings across domains, and critiques evaluation methods. It concludes that the models show strong language and general scientific capabilities but lag expert systems in many conventional NLP tasks, need further multi-step reasoning development, and remain subject to ethical and evaluation concerns.
Problem
Prior assessments cover many tasks and disciplines, but a comprehensive review of ChatGPT and GPT-4 findings is lacking.
Method
The survey reviews assessments of ChatGPT and GPT-4 across language, reasoning, scientific knowledge, and ethics, compares results across tasks, and analyzes evaluation methods.
Results
The survey finds strong language understanding, generation, and general scientific knowledge, but weaker performance than expert systems on many conventional NLP tasks and underdeveloped multi-step reasoning.
Takeaways & Limitations
Evaluation tasks require improved transparency in training corpora and methodology, broader testing domains, and continued attention to fairness, robustness, reliability, and toxicity.
Takeaways & Limitations
GPT models can produce false information when they lack boundaries between positive and negative examples, and human-like fast inference differs from iterative human deliberation.
Abstract
from arXiv · showhide
The emergence of ChatGPT has generated much speculation in the press about its potential to disrupt social and economic systems. Its astonishing language ability has aroused strong curiosity among scholars about its performance in different domains. There have been many studies evaluating the ability of ChatGPT and GPT-4 in different tasks and disciplines. However, a comprehensive review summarizing the collective assessment findings is lacking. The objective of this survey is to thoroughly analyze prior assessments of ChatGPT and GPT-4, focusing on its language and reasoning abilities, scientific knowledge, and ethical considerations. Furthermore, an examination of the existing evaluation methods is conducted, offering several recommendations for future research in evaluating large language models.
1. Introduction
This survey reviews assessments of ChatGPT and GPT-4 across language, scientific knowledge, reasoning, and ethics, addressing the lack of a comprehensive cross-disciplinary synthesis. It finds satisfactory general science performance but weaknesses in multi-step reasoning and evaluation reliability.
- Key findings: ChatGPT performs satisfactorily on general science knowledge and open-response science questions.
- Key findings: ChatGPT can make mistakes on questions requiring multi-step reasoning.
- Evaluation and ethics: Exceptional language proficiency makes factual accuracy difficult for users to assess and raises ethical concerns.
- Evaluation and ethics: Existing evaluations may be unreliable because results depend heavily on prompts and benchmark datasets, including datasets used to train expert systems.
- Scope and contributions: The survey synthesizes assessments of ChatGPT and GPT-4 across language proficiency, scientific knowledge, reasoning, and ethical considerations.It also compares results across tasks and disciplines and critically analyzes evaluation methods.
2. Language and Reasoning Ability
Across language and reasoning tasks, ChatGPT and GPT-4 show strong but uneven abilities: they can surpass some LLM baselines, yet often trail specialized systems and remain vulnerable to hallucination, weak generalization, and reasoning failures.
- Dialogue and generation: ChatGPT surpassed several LLM baselines on dialogue metrics, but SOTA systems still outperformed it on task-oriented and knowledge-grounded dialogue.
- Dialogue and generation: ChatGPT largely underperformed SOTA systems on abstractive and extractive summarization across multiple datasets and languages.A sentence-extraction-then-generation procedure improved faithfulness but still did not reach SOTA performance.
- Dialogue and generation: ChatGPT and GPT-4 underperformed fine-tuned T5 and BART on data-to-text generation, while ChatGPT struggled to produce novel jokes and repeated approximately 90% of generated jokes among 25 jokes.
- Multilingual ability: On Chinese evaluation, GPT-4 scored 76.67 and ChatGPT 66.18, ranking second and third after humans at 96.50.
- Multilingual ability: ChatGPT performed significantly worse than baselines on low- and extremely low-resource languages, despite improving on non-English tasks with English prompts.
- Reasoning: ChatGPT answered 56 of 60 deductive-reasoning questions correctly, compared with 26 of 30 abductive and 33 of 60 inductive questions.
- Reasoning: ChatGPT correctly identified 24 of 30 causes or effects, but causal-identification scores remained below SOTA models.It outperformed relatively weak baselines on causal discovery, while causal-explanation results were inconsistent across automatic metrics.
- Reasoning: Theory-of-mind findings range from human-inferior performance and benchmark contamination concerns to claims that GPT-4 shows advanced ability on broader scenarios.
3. Scientific Knowledge
Assessments across formal and natural sciences show uneven scientific knowledge: GPT models can perform strongly in some exams and domains, but weaknesses remain in mathematical reasoning and specialized knowledge.
- Mathematics: ChatGPT passed only 6 of 17 graduate-level mathematical testing sets, with poor performance on mathematical problem solving.The benchmark covered textbook exercises, Olympiad problems, proof completion, algebra, probability, theorem-proof, and definition understanding.
- Computer Science: GPT-4 scored 24/40 on a computer science exam, slightly exceeding the average score of 23.9 among 200 students.ChatGPT scored 20.5/40, while GPT-4 slightly exceeded the student average; the authors cautioned that online documentation may have supported performance.
- Physics: ChatGPT achieved 53.05% in first-year calculus-based physics, exceeding 90% on clicker and programming questions but performing poorly on homework and exams.Its mathematical difficulties in physics lowered the overall score, which met course-credit requirements but remained below the graduation threshold.
- Medicine: ChatGPT reached college-student level on the USMLE but showed low accuracy in neuro-ophthalmology and high accuracy in general medicine.The contrast indicates uneven performance across medical specialties rather than consistently strong specialized medical knowledge.
- Education: Students perceived ChatGPT responses as linguistically stronger than expert answers despite concerns about scientific accuracy.The study presented 102 physics students with masked ChatGPT and expert responses for evaluation.
- Economics: ChatGPT ranked within the top 9% in microeconomics and top 1% in macroeconomics among college students, answering 19/30 and 26/30 questions correctly.These results came from the Test of Understanding of College Economics, using comparisons with 3,255 microeconomics and 2,789 macroeconomics students.
4. Ethical Considerations
Ethical assessments identify progress in mitigating some social biases and toxic outputs, but also show persistent vulnerabilities to language disparities and role-playing prompt injections.
- Fairness: ChatGPT responses were much worse in languages other than English, indicating a fairness disparity across languages.Fairness assessments considered gender, race, language, and culture; other studies examined gender and race bias separately.
- Toxicity: Only 0.5% of ChatGPT responses were toxic under toxic prompting, but role-playing could still induce offensive content.The low toxicity rate was not robust, and prompt injections achieved through role-playing remained effective.
5. Discussion
The discussion portrays GPT models as capable but unlike humans in reasoning, knowledge representation, and concept use, while emphasizing hallucinations, evaluation instability, and ethical risks.
- Comparing GPT versus Humans: GPT models may match or surpass average human accuracy in some exams without demonstrating human-like performance or intelligence across individual cases.Average scores can conceal specialist strengths alongside failures on examples that humans find relatively easy.
- Comparing GPT versus Humans: GPT hallucinations may arise because next-word pre-training emphasizes what is right without adequately learning what is wrong.The paper uses penguin flight as an example of how learned positive associations can produce false information when boundaries are unclear.
- Comparing GPT versus Humans: Humans can deliberate iteratively and backtrack, whereas GPT models mainly use feedforward inference and linear chains of thought.“Let’s think step by step” can support multi-step solutions, but the paper describes planning and alternative exploration as severe limitations.
- Comparing GPT versus Humans: GPT knowledge is sensitive to wording and language, so semantically equivalent questions may elicit different or factually incorrect responses.The paper links this context dependence to entangled internal representations and notes that more training cannot remove the fundamental dependence on textual context.
- Evaluation: Evaluation results are difficult to compare because ChatGPT changes over time, prompts influence outcomes, and assessment data may leak into later model versions.The discussion recommends transparent prompt design and fair comparisons with baselines while retaining established concerns about corpora, metrics, and human evaluation.
- Ethics: Human concept mappings were more diverse than ChatGPT’s, whose outputs activated fewer concepts across the mapped space.MetaPro was used to compare target and source concepts from parallel human and ChatGPT answers; bright dots indicate activated concepts and grey dots unactivated concepts.
- Ethics: RLHF can mitigate biased, inaccurate, and toxic responses, but human-biased feedback and opaque training data create additional bias, privacy, and leakage concerns.The discussion identifies system gaming, positive reward cycles, tangled social norms, and possible use of user inputs for fine-tuning as risks.
6. Conclusion and Recommendations
The survey finds strong language and general scientific capabilities alongside weaknesses in expert-level NLP, multi-step reasoning, ethics, and evaluation reliability. It recommends broader, task-agnostic evaluation, continued fundamental research, and evolving regulation for AI-generated content.
- The surveyed GPT models show strong language understanding, generation, and general scientific knowledge, but lag behind expert systems on many conventional NLP tasks.
- Multi-step reasoning remains an area where the models need further development, while fairness, robustness, reliability, toxicity, and domain-specific ethics remain concerns.
- Task-agnostic evaluation is recommended because benchmark- and task-specific assessments risk data contamination during pre-training.
- Indirect concept-mapping approaches are presented as a promising model for evaluating nonopen-ended capabilities through responses that reflect cognition and emotional states.
- Fundamental computational-linguistics research remains valuable for understanding human language and because expert models still lead in many resource-rich areas.
- The paper argues that complex tasks should be decomposed into subtasks and that regulation should evolve quickly to address misuse of AI-generated content.
- Because AI-generated content can be produced faster than humans and may dilute diversity of human thought, proper controls are needed.
A. ChatGPT and GPT-4 benchmark
The survey benchmarks ChatGPT and GPT-4 across tasks by comparing them with baselines, humans, or ground truth. Its visualizations encode the size and direction of performance gaps across evaluation settings.
- Tables 2–7 visualize performance gaps between surveyed GPT models and baselines across different tasks.
- The discrepancy measure is g/b −1, comparing a GPT model’s average major-metric score g with a baseline’s average score b.
- Red marks GPT models outperforming baselines, blue marks lagging performance, teal marks gaps from ground truth, and gray marks missing comparisons.
- Table 1 defines task abbreviations, while Table 2 reports linguistic and reasoning performance against baseline models.
- Table 3 compares ChatGPT and GPT-4 on linguistic and reasoning tasks with humans or ground truth.
ChatGPT GPT-4 ChatGPT
Table 4 organizes multilingual benchmark results by language-resource level and compares ChatGPT and GPT-4 with baseline models.
- Table 4 reports ChatGPT and GPT-4 performance on multilingual tasks against baseline models.
- Languages are grouped into high-, medium-, low-, and extremely-low-resource categories.
- The listed language groups are en2zh-vi, tr-hi, bn-kn, and swas, respectively.
ChatGPT G4
The supplied table caption identifies a multilingual comparison of ChatGPT and GPT-4 with humans or ground truth, but provides no readable task-level findings.
- Table 5 compares ChatGPT and GPT-4 on multilingual tasks with humans or ground truth.
- The supplied table text does not state which multilingual tasks or language groups are included.
- The supplied table text does not report any performance outcomes from the comparison.
ChatGPT GPT-4
The surveyed material includes tables comparing ChatGPT and GPT-4 performance on scientific knowledge with baselines, humans, or ground truth, alongside summaries of surveyed works.
- Table 6 compares ChatGPT and GPT-4 performance on scientific knowledge with baselines.
- Table 7 compares ChatGPT and GPT-4 performance on scientific knowledge with humans or ground truth.
- Tables 8 and 9 summarize the surveyed works, with Table 8 defining CG as ChatGPT and G4 as GPT-4.