Source-linked AI summary
ChatGPT: The End of Online Exam Integrity?
Teo Susnjak
TL;DR
The paper investigates whether ChatGPT can perform complex reasoning and generate human-like answers relevant to online examinations. Through cross-disciplinary question generation, answering, and self-critique, it finds strong critical-thinking and realistic-text capabilities that threaten online exam integrity, while noting important evaluation and prevention limits.
Problem
Online exams face academic-integrity risks, and the implications of ChatGPT’s ability to answer demanding questions and produce human-like text require investigation.
Method
The study tests ChatGPT across four disciplines by having it generate undergraduate critical-thinking questions, answer them, and critically evaluate its responses.
Results
ChatGPT produced clear, precise, relevant, deep, broad, logically coherent responses and demonstrated critical reasoning and human-like prose.
Takeaways & Limitations
ChatGPT’s capabilities create a significant threat to online exam integrity, especially in tertiary education where online examinations are increasingly common.
Takeaways & Limitations
The investigation was preliminary and could be strengthened by independent subject experts and prior examination questions from multiple courses.
Abstract
from arXiv · showhide
This study evaluated the ability of ChatGPT, a recently developed artificial intelligence (AI) agent, to perform high-level cognitive tasks and produce text that is indistinguishable from human-generated text. This capacity raises concerns about the potential use of ChatGPT as a tool for academic misconduct in online exams. The study found that ChatGPT is capable of exhibiting critical thinking skills and generating highly realistic text with minimal input, making it a potential threat to the integrity of online exams, particularly in tertiary education settings where such exams are becoming more prevalent. Returning to invigilated and oral exams could form part of the solution, while using advanced proctoring techniques and AI-text output detectors may be effective in addressing this issue, they are not likely to be foolproof solutions. Further research is needed to fully understand the implications of large language models like ChatGPT and to devise strategies for combating the risk of cheating using these tools. It is crucial for educators and institutions to be aware of the possibility of ChatGPT being used for cheating and to investigate measures to address it in order to maintain the fairness and validity of online exams for all students.
1 Introduction
The expansion of online higher education has intensified concerns about exam cheating, while ChatGPT introduces a new risk by answering demanding university-level questions with coherent, realistic text. The study therefore examines ChatGPT’s reasoning capabilities and implications for academic integrity.
- 1 Introduction: Online learning and examinations have expanded, with this shift accelerated by the COVID-19 pandemic and unlikely to reverse soon.Remote learning benefits have been increasingly recognized by institutions and students.
- 1 Introduction: Online exams create cheating risks through anonymity, limited direct supervision, and easy access to shared resources.
- 1 Introduction: Research has not definitively quantified dishonest practices in online assessments, although several studies indicate cheating is prevalent and may exceed face-to-face rates.
- 1 Introduction: Existing responses include redesigned assessments, proctoring, plagiarism detection, security measures, policy revisions, and educational campaigns, but their overall effectiveness remains insufficiently established.
- 1 Introduction: ChatGPT can generate accurate answers to difficult questions requiring analysis, synthesis, and application, creating a threat even for exams designed to assess higher-order reasoning.
- 1 Introduction: The study analyzes ChatGPT’s ability to answer coherent, non-trivial university questions across disciplines and raises an urgent warning about existing integrity safeguards.
2 Background
Prior research identifies persistent cheating and unresolved validity, privacy, bias, and reliability problems in online assessment and proctoring. The background literature also describes alternative assessment strategies, while emphasizing the need for further research.
- 2 Background: The datasets used to train ChatGPT had not been publicly released, limiting transparency about its training data.
- 2 Background: Online examinations raise challenges involving cheating, technology access, and the absence of standardized approaches, motivating calls for fair, valid, and reliable designs.
- 2 Background: Proctoring technologies raise concerns about privacy, bias, software validity, reliability, intrusive monitoring, and unequal effects related to bandwidth.
- 2 Background: Systematic reviews report that cheating is a significant concern and more prevalent online than in traditional face-to-face examinations.
- 2 Background: Cheating persists in both online and invigilated paper examinations, while evidence about the effects of invigilation and online security remains conflicting.
- 2 Background: Educators have considered replacing multiple-choice questions with short-answer or critical-thinking questions and using tighter time limits, although students may perceive such exams as harder.
3 Methodology
The study tested ChatGPT across multiple undergraduate disciplines by having it generate difficult critical-thinking questions, answer them, and critically evaluate its responses. Responses were assessed using universal intellectual standards and additional criteria including originality and persuasiveness.
- 3 Methodology: The methodology comprised question generation, answer generation, and critical evaluation of the generated answers.
- 3 Methodology: The study used ChatGPT’s publicly accessible online portal for the experiment.
- 3 Methodology: ChatGPT generated challenging, scenario-based questions intended for undergraduate students across Machine Learning, Marketing, History, and Education.
- 3 Methodology: Answers were requested in several paragraphs using 500 words, examples, and supporting arguments.
- 3 Methodology: ChatGPT was prompted to identify strengths, weaknesses, and improvement suggestions in its own answers.
- 3.1 Evaluation of responses: Response evaluation used universal intellectual standards covering relevance, clarity, accuracy, precision, depth, breadth, logic, persuasiveness, and originality.
- 3.1 Evaluation of responses: The evaluation additionally considered whether responses offered new insights rather than repeating established information.
4 Results
ChatGPT produced clear, coherent, relevant, and logically structured responses across Education, Machine Learning, History, and Marketing tasks. The responses also showed depth, breadth, persuasiveness, and specific examples, although accuracy assessment was limited outside Machine Learning.
- Response quality: ChatGPT responses across four disciplines were clear, coherent, and appropriately structured for their intended audiences.The responses used straightforward language, appropriate technical vocabulary, natural-language conventions, and an intentional flow of ideas.
- Accuracy and originality: Accuracy was directly attested only for the Machine Learning material, while expert assessment in Marketing, Education, and History was outside the study’s scope.The author described the Machine Learning account of overfitting and its mitigation techniques as accurate, but did not obtain subject-expert evaluations for the other disciplines.
- Response quality: The responses were relevant to difficult prompts requiring hypothetical questions, answers, and critical analyses.All responses were reported as on-topic for both the disciplinary subject matter and the intent of the requests.
- Critical-thinking performance: ChatGPT demonstrated depth through complex questions, supporting rationales, substantial critiques, and suggested improvements.Answers were constrained to 500 words yet included strategies and examples across all four disciplines.
- Critical-thinking performance: The answers demonstrated breadth by explaining two scenarios in each case and adding further examples through improvement suggestions.The breadth was observed within the constraints imposed on the responses.
- Critical-thinking performance: Arguments and evidence were presented clearly and logically, with efforts to address potential counterarguments, although persuasiveness varied by reader perspective.Responses were expressed confidently, without reservations, regardless of whether the claims were correct.
5 Discussion
The study argues that ChatGPT can perform critical reasoning and produce human-like prose, creating immediate challenges for online examination integrity. Proposed safeguards include multimodal or oral assessment, proctoring, and detection tools, but their effectiveness remains constrained and requires continued evaluation.
- 5 Discussion: ChatGPT demonstrated critical thinking beyond information retrieval, producing clear, precise, relevant, logically coherent responses across constrained tasks.The study describes its responses as sufficiently deep and broad, with clear exposition and appropriate examples.
- 5 Discussion: ChatGPT can critique its own responses, discuss strengths and weaknesses, suggest improvements, and express ideas in prose that appears comparable to human writing.These behaviors are presented as evidence of conceptualization and higher-order thinking rather than mere memorization.
- 5 Discussion: Human-like generated responses raise serious questions about the reliability and validity of online exams and the potential for cheating.The concern is especially relevant as online examination integrity faces immediate consequences from these capabilities.
- 5.1 Recommendations for mitigating strategies: Multimodal questions, including images and recorded video, could exploit ChatGPT’s current text-only input limitation.The paper notes that this limitation may make accurate response generation more difficult, while also warning that future systems may incorporate images, video, and audio.
- 5.1 Recommendations for mitigating strategies: Oral exams, proctoring, secure exam technologies, and GPT-output detectors are proposed safeguards, but none is presented as a foolproof solution.The discussion also questions plagiarism detection, notes that detection systems may be costly and require further research, and reports that ChatGPT cannot reliably verify text generated in prior sessions.
- 5.2 Limitations: The investigation is preliminary, and the authors recommend future evaluation by independent subject experts using experts and examination questions from varied courses.The study also reports that ChatGPT is capable of generating effective questions.
6 Conclusion
The study examined whether ChatGPT can perform high-order thinking tasks and generate human-like text that could facilitate cheating in online examinations. It found that ChatGPT poses a significant threat to online-exam integrity and that proposed countermeasures are not foolproof.
- ChatGPT was evaluated on high-order thinking and human-like text generation relevant to potential academic dishonesty in online examinations.The agent was prompted to generate questions, provide rationales and answers, and produce critiques.
- ChatGPT presents a significant threat to online-exam integrity, especially in tertiary education where online exams are increasingly common.
- ChatGPT demonstrates critical thinking and generates highly realistic text with little input, making exam cheating possible.
- Invigilated and oral exams and advanced proctoring may help combat the threat, but they are not perfect solutions.
- Further research is needed on AI-text detection and strategies for addressing cheating with large language models.
A Examples of Multiple-choice questions and answers generated by ChatGPT
This section presents multiple-choice questions generated by ChatGPT across four disciplines, together with answers and associated explanations.
- ChatGPT generated multiple-choice questions across four disciplines for the study.
- The generated examples include answer choices for each question.
- The section provides answers with associated explanations for the generated questions.
A.1 Machine Learning
The machine-learning example asks which statement is not a disadvantage of SVM classification. The provided answer is feature-scaling insensitivity, while the explanation identifies several other disadvantages.
- The SVM question asks which option is not a disadvantage of using an SVM model for classification.
- The listed SVM disadvantages include kernel and hyperparameter sensitivity, poor generalization on non-linearly separable data, slow training, and possible overfitting.
- The provided answer is C: SVMs are not sensitive to the scaling of the input features.
- The explanation describes SVMs as linear classifiers seeking a hyperplane that maximally separates classes.
A.2 Education
The education example asks which theory best explains learning through observation and imitation. It identifies Bandura’s social learning theory and explains learning through watching and mimicking others.
- The question asks which theory best explains learning through observation and imitation.
- The answer given is C: Bandura’s social learning theory.
- Bandura’s theory explains that individuals learn by watching and mimicking others’ actions, through vicarious learning or modeling.
- The explanation contrasts Bandura’s theory with Piaget’s, Vygotsky’s, and Bloom’s frameworks.
A.3 Marketing
Lifestyle branding persuades consumers through emotional appeals and aspirational messaging by presenting products as enhancing a desired lifestyle or image.
- Psychological pricing is not identified as involving emotional appeals or aspirational messaging.
- Lifestyle branding creates an emotional connection by presenting products as a way to enhance consumers’ desired lifestyle or image.
- Aspirational messaging and emotional appeals persuade consumers to purchase the product.
A.4 History
The Indian Mutiny of 1857 was a significant rebellion against British East India Company rule in India. Its stated significance was the abolition of the Company and transfer of power to the British Crown.
- The Indian Mutiny of 1857 resulted in the abolition of the East India Company and transfer of power to the British Crown.
- The mutiny was a widespread rebellion against the British East India Company, which governed India at the time.
- The uprising began as a protest against animal fat used to grease rifle cartridges, offending Hindus and Muslims.
- The cartridge controversy quickly escalated into a broader uprising against British rule.