Source-linked AI summary

Perception, performance, and detectability of conversational artificial intelligence across 32 university courses

Hazem Ibrahim, Fengyuan Liu, Rohail Asim, Balaraju Battu, Sidahmed Benabderrahmane, Bashar Alhafni, Wifag Adnan, Tuka Alhanai, Bedoor AlShebli, Riyadh Baghdadi, Jocelyn J. Bélanger, Elena Beretta, Kemal Celik, Moumena Chaqfeh, Mohammed F. Daqaq, Zaynab El Bernoussi, Daryl Fougnie, Borja Garcia de Soto, Alberto Gandolfi, Andras Gyorgy, Nizar Habash, J. Andrew Harris, Aaron Kaufman, Lefteris Kirousis, Korhan Kocak, Kangsan Lee, Seungah S. Lee, Samreen Malik, Michail Maniatakos, David Melcher, Azzam Mourad, Minsu Park, Mahmoud Rasras, Alicja Reuben, Dania Zantout, Nancy W. Gleason, Kinga Makovi, Talal Rahwan, Yasir Zaki

arXiv:2305.13934v1cs.CYcs.AI

TL;DR

The paper addresses limited evidence about ChatGPT’s university-level performance, detectability, and educational acceptance. It compares ChatGPT with students, evaluates detection and obfuscation, and surveys students and educators. ChatGPT matched or exceeded students on multiple courses, detection was unreliable, and students and educators showed emerging opposing norms around its use.

  • Problem

    Evidence was limited about ChatGPT’s performance versus university students, detectability in school work, and students’ and educators’ perspectives on its use.

  • Method

    The study compares ChatGPT and student answers across 32 courses, tests detection algorithms and obfuscation, and surveys participants across five countries and one institution.

  • Results

    ChatGPT’s performance was comparable or superior to students’ on nine of 32 courses, while detection algorithms misclassified answers and were defeated by obfuscation.

  • Takeaways & Limitations

    Students showed support for using ChatGPT in assignments, while professors showed support for treating its use as plagiarism.

  • Takeaways & Limitations

    Whether policies such as banning ChatGPT can be effectively enforced remains unknown.

Abstract

from arXiv · show

The emergence of large language models has led to the development of powerful tools such as ChatGPT that can produce text indistinguishable from human-generated work. With the increasing accessibility of such technology, students across the globe may utilize it to help with their school work -- a possibility that has sparked discussions on the integrity of student evaluations in the age of artificial intelligence (AI). To date, it is unclear how such tools perform compared to students on university-level courses. Further, students' perspectives regarding the use of such tools, and educators' perspectives on treating their use as plagiarism, remain unknown. Here, we compare the performance of ChatGPT against students on 32 university-level courses. We also assess the degree to which its use can be detected by two classifiers designed specifically for this purpose. Additionally, we conduct a survey across five countries, as well as a more in-depth survey at the authors' institution, to discern students' and educators' perceptions of ChatGPT's use. We find that ChatGPT's performance is comparable, if not superior, to that of students in many courses. Moreover, current AI-text classifiers cannot reliably detect ChatGPT's use in school work, due to their propensity to classify human-written answers as AI-generated, as well as the ease with which AI-generated text can be edited to evade detection. Finally, we find an emerging consensus among students to use the tool, and among educators to treat this as plagiarism. Our findings offer insights that could guide policy discussions addressing the integration of AI into educational frameworks.

Significance statement

The study addresses the challenge of integrating artificial intelligence into educational frameworks. It fills gaps in evidence about ChatGPT’s university-level performance, detectability, and educational perceptions.

  • The literature lacked a systematic evaluation of generative AI tools on university-level courses.
  • The study examines both the detectability of AI-generated work and students’ and educators’ perspectives on its educational use.
  • Its findings provide insights for policy discussions about student evaluation frameworks in the age of artificial intelligence.

Introduction

Generative AI can create new content, and ChatGPT produces human-like conversational responses across many prompts. In education, its use raises academic-integrity concerns while systematic evidence about performance and detectability remains limited.

  • Generative AI uses machine-learning algorithms to build on existing material and create new content.
  • ChatGPT generates human-like textual responses across many languages through an ongoing dialogue with users.
  • ChatGPT’s ability to write essays and solve assignments has intensified concerns about academic-integrity violations by students.
  • The study compares ChatGPT with students across 32 university-level courses, evaluates detection algorithms and obfuscation, and surveys participants in five countries.
  • ChatGPT’s performance was comparable or superior to students’ on nine of the 32 courses, while current detection algorithms produced both types of misclassification.

Results

Across university courses and surveys, ChatGPT performed comparably to students in many settings, while perceptions of its educational use varied between students and educators. Detection systems were vulnerable to both false positives and simple text obfuscation.

  • Performance: ChatGPT received an average grade of 7.5 on creativity questions, compared with 7.9 for students.
  • Performance: Math-related and trick questions showed the largest performance gaps, with humans outperforming ChatGPT in these areas.
  • Global survey: 74% of surveyed students said they would use ChatGPT, mainly to improve skills and save time.
  • Global survey: Students who would not use ChatGPT cited not knowing how or having no need more often than fear of penalties or unethical conduct.
  • NYUAD survey: At NYUAD, 57% of students planned to use ChatGPT, while 69% of professors planned to treat its use as plagiarism.
  • Detectability: Quillbot increased false-negative rates from 49% to 98% for OpenAI’s classifier and from 32% to 95% for GPTZero.

Discussion

ChatGPT performed comparably or better than students in many university courses, while students and educators expressed conflicting norms about its use and detection. These findings highlight challenges for existing student-evaluation frameworks.

  • Perceptions: Students and educators agreed that ChatGPT use in school work should be acknowledged, and that it could increase competitiveness for non-native English speakers.Students also expected to outsource mundane future-job tasks to ChatGPT and focus on substantive and creative work.
  • Norms and policy: Students generally planned to use ChatGPT and believed peers would approve, while professors planned to treat its use as plagiarism and expected peer approval.These opposing expectations imply a conflict between student-use norms and faculty-enforcement norms.
  • Performance and detection: 38% of courses showed ChatGPT performing comparably or better than students, covering 12 of 32 courses.The study presents this comparison as evidence that AI-generated work can materially affect student evaluation.
  • Performance and detection: Figure 4 evaluates GPTZero and an AI text classifier using confusion matrices, a Quillbot obfuscation attack, and comparisons across courses.The cited discussion identifies ease of evading current classifiers as a challenge for enforcement.

Methods

The study combines course-based comparisons of student and ChatGPT answers with surveys of students and educators across five countries and at NYUAD. It also tests AI-text classifiers and their susceptibility to Quillbot-based obfuscation.

  • Course performance analysis: Faculty selected 10 text-based questions from courses and three student submissions per question, while ChatGPT generated three answers to each question.Questions came from labs, homework, assignments, quizzes, or exams, and submissions were randomized for grading.
  • Course performance analysis: Questions could include tables, programming code, mathematics, and fill-the-gap formats, but excluded multiple-choice questions, images, diagrams, and attachments.The guidelines constrained the assessment format while allowing several question types.
  • Course performance analysis: Faculty classified each question by knowledge and cognitive-process dimensions and recorded whether it involved mathematics, code, specific sources or methods, or trick-question characteristics.These annotations supported analyses of how ChatGPT performance varied across question properties.
  • Perception surveys: Survey participants came from Brazil, India, Japan, the United Kingdom, and the United States, with at least 200 students and 100 educators recruited per country.The survey was piloted in the United States before the cross-country study.
  • Perception surveys: The NYUAD survey included 151 students and 60 professors and added demographic, academic, and normative-expectation questions.Student and faculty respondents were asked about expectations surrounding ChatGPT use and its treatment as plagiarism.
  • Detection analysis: GPTZero and OpenAI’s AI text classifier were applied to student and ChatGPT submissions to assess their classification of AI-generated text.The classifiers were designed specifically to determine whether text was generated using AI.
  • Detection analysis: The Quillbot attack maximized synonym substitutions and ran ChatGPT submissions through each available mode before re-analyzing the outputs with both classifiers.The procedure tested whether readily available paraphrasing could alter classifier predictions.

Ethics statement

The study received institutional ethics approval and followed relevant guidelines, with informed consent obtained from all participants.

  • Ethics statement: The study was approved by NYU Abu Dhabi’s Institutional Review Board under protocol HRPP-2023-5.The researchers state that all research followed relevant guidelines and regulations.
  • Ethics statement: Informed consent was obtained from participants in every segment of the study.
Loading 2305.13934v1…