Source-linked AI summary

Will ChatGPT get you caught? Rethinking of Plagiarism Detection

Mohammad Khalil, Erkan Er

arXiv:2302.04335v1cs.AI

TL;DR

The paper addresses concerns that ChatGPT can generate academic essays that evade plagiarism detection. It evaluates 50 ChatGPT-generated essays with Turnitin and iThenticate, while also asking ChatGPT to verify authorship; most essays were judged highly original by conventional tools, whereas ChatGPT identified more generated essays.

  • Problem

    ChatGPT’s growing use for academic essay generation raises concerns about whether conventional plagiarism tools can detect AI-produced work.

  • Method

    The study descriptively analyzes originality scores for 50 ChatGPT-generated essays using Turnitin and iThenticate, supplemented by ChatGPT-based authorship verification.

  • Results

    40 of 50 essays had similarity scores of 20% or less, while ChatGPT correctly predicted 46 essays as generated by ChatGPT compared with 10 identified by plagiarism-detection tools at critical similarity levels.

  • Takeaways & Limitations

    The findings suggest that conventional plagiarism detection may need reconsideration as AI-generated academic content becomes more capable of avoiding similarity-based detection.

  • Takeaways & Limitations

    The study examines only ChatGPT, depends on Turnitin and iThenticate accuracy, and uses a sample of 50 essays that may be insufficient for generalization.

Abstract

from arXiv · show

The rise of Artificial Intelligence (AI) technology and its impact on education has been a topic of growing concern in recent years. The new generation AI systems such as chatbots have become more accessible on the Internet and stronger in terms of capabilities. The use of chatbots, particularly ChatGPT, for generating academic essays at schools and colleges has sparked fears among scholars. This study aims to explore the originality of contents produced by one of the most popular AI chatbots, ChatGPT. To this end, two popular plagiarism detection tools were used to evaluate the originality of 50 essays generated by ChatGPT on various topics. Our results manifest that ChatGPT has a great potential to generate sophisticated text outputs without being well caught by the plagiarism check software. In other words, ChatGPT can create content on many topics with high originality as if they were written by someone. These findings align with the recent concerns about students using chatbots for an easy shortcut to success with minimal or no effort. Moreover, ChatGPT was asked to verify if the essays were generated by itself, as an additional measure of plagiarism check, and it showed superior performance compared to the traditional plagiarism-detection tools. The paper discusses the need for institutions to consider appropriate measures to mitigate potential plagiarism issues and advise on the ongoing debate surrounding the impact of AI technology on education. Further implications are discussed in the paper.

1 Introduction

ChatGPT’s accessibility and capabilities have intensified concerns about plagiarism and academic honesty in education. This study examines those concerns by testing whether ChatGPT-generated essays can avoid detection by plagiarism software.

  • Chatbots use natural language processing and machine learning to simulate human-like conversations across platforms, including education.
  • ChatGPT is an accessible deep-learning chatbot designed to simulate conversation with Internet users and handle diverse information tasks.
  • ChatGPT’s ability to generate high-quality scholarly text has raised concerns about plagiarism in educational settings.
  • Some schools and school districts have prohibited ChatGPT on student devices and networks in response to these concerns.
  • The study asks ChatGPT 50 open-ended questions and checks the resulting essays with iThenticate and Turnitin to provide empirical evidence about plagiarism avoidance.

2 Background

The background describes chatbots as increasingly important educational tools while placing ChatGPT within wider concerns about cheating, proctoring, and plagiarism. It also introduces the detection services used to identify copied work.

  • Chatbots simulate human-like conversations through text or audio and have become increasingly popular since modern systems emerged around 2016.
  • Chatbots in education support skill improvement, task assistance, content teaching, and efficiency through task automation.
  • ChatGPT is described as capable of creating code, performing complex mathematical operations, and producing essays, stories, and poems.
  • Remote assignments and tests have expanded online proctoring, while commercial educational infrastructure raises concerns about trustworthiness and academic cheating.
  • Plagiarism means presenting someone else’s work or ideas as one’s own without proper attribution, including copied text or images.
  • Turnitin and iThenticate are widely used anti-plagiarism services that detect copied work by comparing submissions with extensive content databases.

3 Methodology

The study quantitatively evaluates originality in ChatGPT-generated essays using two plagiarism-detection services and an additional ChatGPT-based verification step. Essays are generated from 50 topics, checked against large databases, and analyzed descriptively.

  • The study uses a descriptive quantitative design to analyze originality scores for AI-generated content.
  • The authors proposed 50 topics and instructed ChatGPT to write a 500-word essay for each, saving outputs as separate student-submission files.
  • The 50 essays were split between Turnitin (n= 25) and iThenticate (n= 25) for comparison against Internet articles, academic papers, and webpages.
  • Fig. 2 presents an iThenticate Similarity score for (n= 7) essays, representing the plagiarism proportion produced by the software.
  • ChatGPT additionally inspected all 50 essays to identify whether they were generated by itself, providing a comparison with conventional detection methods.
  • The resulting plagiarism measures were analyzed descriptively using quantitative originality scores.

4 Findings

Across 50 ChatGPT-generated essays, plagiarism similarity varied substantially across examples and detection tools. ChatGPT’s reverse-engineering check identified most essays as generated by itself.

  • Similarity scores ranged from 0% to 64% across the essays, with examples scoring 5%, 14%, and 64%.The examples concerned essays on Robots, Learning theories, and Laws of physics.
  • 17 of 25 essays assessed by iThenticate had similarity below 10%, while the average similarity score was 8.76.Three essays scored 20–40%, and none exceeded 40%.
  • Turnitin results showed higher similarity overall, with an average score of 13.72 compared with 8.76 for the initial result set.One essay exceeded 40% similarity, and six essays had scores between 20% and 40%.
  • 4.1 Reverse engineering: ChatGPT identified 46 of 50 essays as generated by itself, leaving four undetected in the reverse-engineering check.The reported accuracy was over 92%.

5 Discussion

The study finds that ChatGPT-generated essays often evade conventional plagiarism detection, while ChatGPT’s self-identification performed better. However, the findings are constrained by the study’s design and sample.

  • Implications: The results suggest that students could submit ChatGPT-produced essays while avoiding detection, particularly when assignments require personal reflection or interpretation.Interpretative examples included cultural differences, good-teacher characteristics, and leadership.
  • Alternative detection: ChatGPT’s self-verification correctly identified 46 generated essays, whereas plagiarism-detection tools identified only 10 essays with critical similarity.The authors therefore question whether similarity checking alone is adequate for AI-generated content.
  • Limitations: The study is limited to ChatGPT, depends on Turnitin and iThenticate accuracy, uses 50 essays, and includes texts that sometimes fell below 500 words.The authors also state that ChatGPT-based reverse engineering remains unverified against similarity scores.

6 Conclusions

The paper recognizes educational benefits from large language models while warning that ChatGPT can provide essays on demand and be used unethically. It recommends guidance, active learning, academic-integrity education, and institutional policies.

  • 6 Conclusions: Large language models such as ChatGPT and Google Bard can improve students’ educational experience and facilitate teachers’ tasks, but may also be used unethically.The paper specifically warns that students may use chatbots to produce academic essays on demand.
  • 6 Conclusions: Teachers are advised to assign work beyond basic tasks, foster active engagement and critical thinking, explain ChatGPT’s limitations, and clarify academic-integrity expectations.The recommendations also call for clear syllabus guidelines on AI use.
  • 6 Conclusions: Students are advised to use ChatGPT to improve competencies and learning, not as a substitute for original thinking and writing.They should also understand proper and ethical use in their courses and the consequences of relying solely on the tool.
  • 6 Conclusions: Institutions are advised to understand large language models’ educational potential, maintain stakeholder communication, and implement clear AI-use policies and guidelines.The stated stakeholders include researchers and IT support.
  • 6 Conclusions: The paper recommends training students, faculty, and staff on academic integrity and responsible use of AI tools in education.
Loading 2302.04335v1…