Source-linked AI summary
Game of Tones: Faculty detection of GPT-4 generated content in university assessments
Mike Perkins, Jasper Roe, Darius Postma, James McGaughran, Don Hickerson
TL;DR
This study examines whether university assessments can withstand GPT-4-generated submissions and whether academic staff can identify them using Turnitin AI detection. In a School of Business assessment experiment, GPT-4 submissions were marked by faculty, with detection and grading results showing substantial challenges for academic integrity.
Problem
Advances in large language models make human and AI-generated academic content difficult to distinguish, while evidence on academic staff detecting GPT-4 content was lacking.
Method
Researchers created 22 GPT-4 responses to eligible School of Business assessments and incorporated them into assessment processes reviewed by 15 academic staff using Turnitin Feedback Studio and AI detection.
Results
54.8% was the mean percentage of AI content detected, although 91% of submissions contained some highlighted AI content; faculty identified 54.5% of submissions as potentially AI-generated.
Takeaways & Limitations
The findings support improving AI detection tools, increasing faculty training and awareness, and developing assessment strategies that incorporate or resist AI use.
Takeaways & Limitations
The study was conducted in one Business School with a small group of academic staff, limiting how widely its findings may apply.
Abstract
from arXiv · showhide
This study explores the robustness of university assessments against the use of Open AI's Generative Pre-Trained Transformer 4 (GPT-4) generated content and evaluates the ability of academic staff to detect its use when supported by the Turnitin Artificial Intelligence (AI) detection tool. The research involved twenty-two GPT-4 generated submissions being created and included in the assessment process to be marked by fifteen different faculty members. The study reveals that although the detection tool identified 91% of the experimental submissions as containing some AI-generated content, the total detected content was only 54.8%. This suggests that the use of adversarial techniques regarding prompt engineering is an effective method in evading AI detection tools and highlights that improvements to AI detection software are needed. Using the Turnitin AI detect tool, faculty reported 54.5% of the experimental submissions to the academic misconduct process, suggesting the need for increased awareness and training into these tools. Genuine submissions received a mean score of 54.4, whereas AI-generated content scored 52.3, indicating the comparable performance of GPT-4 in real-life situations. Recommendations include adjusting assessment strategies to make them more resistant to the use of AI tools, using AI-inclusive assessment where possible, and providing comprehensive training programs for faculty and students. This research contributes to understanding the relationship between AI-generated content and academic assessment, urging further investigation to preserve academic integrity.
Introduction
The introduction frames GPT-4 and other generative AI tools as creating opportunities alongside risks for academic assessment and integrity. It defines the study’s objectives as evaluating GPT-4-generated work, Turnitin detection, faculty judgments, and implications for assessment design.
- GPT-4 and other generative AI tools can produce fluent, detailed, natural-sounding text that is difficult to distinguish from human-written work.
- AI-generated assessed work raises concerns about authorship, authenticity, academic integrity, and the educational value of assignments.
- The study evaluates GPT-4 paper quality, perceived authenticity, Turnitin’s detection effectiveness, faculty judgments, and staff perceptions of AI-generated content.
- Assessment strategies are presented as needing adaptation to maintain academic rigor while allowing responsible use of AI and some cognitive offloading.
- The paper positions AI detection as a basis for educating students, training staff, and adapting assessment methods rather than relying on a detect-and-punish approach.
Literature Review
Prior research reports difficulty distinguishing AI-generated from human-written work and limited detection-tool effectiveness, while lacking empirical studies of GPT-4 in realistic university assessment settings. This study addresses that gap.
- AI-generated text can be difficult for educators to distinguish from human-written work, challenging academic standards and fair assessment.
- Existing detection tools have limitations, particularly when users deliberately alter or paraphrase AI-generated content.
- The literature lacks empirical studies evaluating academic staff detection of GPT-4 content and AI detectors under realistic student-assessment conditions.
- Earlier studies found limited human detection accuracy, including 59.5% correct identification of LLM-generated samples for GPT-2.
- Commercial detectors report widely varying claimed performance, including OpenAI’s 26% AI-text identification rate and Originality.AI’s 95.93% true-positive rate.
- The study addresses this gap because evolving LLM capabilities raise concerns about fairness and equality in higher-education assessment.
Methodology
The study created GPT-4 submissions for real university assessments and had experienced faculty mark them using Turnitin workflows. Prompt engineering was used to satisfy task requirements and reduce detection likelihood.
- Context: The study was conducted in one Southeast Asian university’s Business School, where fifteen experienced faculty members used Turnitin Feedback Studio for blind marking.
- Experimental design: The research team reviewed 70 assessments, retained 25 eligible assignments, and generated 22 GPT-4 responses for submission through a test student profile.
- Experimental design: After marking, researchers recorded grades, Turnitin AI scores, misconduct reports, and other relevant information before debriefing and training faculty.
- Experimental design: The generated assessments included essays, reports, and case analyses across topics such as marketing, consumer behaviour, and supply-chain management.
- Experimental design: Prompt engineering targeted assessment requirements and used adversarial techniques, including requests for language complexity associated with non-native English speakers.
- Experimental design: The team did not use paraphrasing tools and made no substantive manual adjustments beyond basic formatting.
Results
Across 22 GPT-4 submissions, faculty reported 12 as potentially AI-generated, while Turnitin detected some AI content in 91% but identified only 54.8% of content on average. AI-generated work received grades close to genuine work, and marker feedback varied substantially.
- Summary results: 12 (54.5%) of 22 GPT-4 submissions were identified by academic staff as potentially AI-generated.
- Summary results: 91% of submissions contained some content highlighted by Turnitin AI detect, but the mean percentage of AI content detected was 54.8%.
- Summary results: Genuine submissions received a mean score of 54.4, whereas AI-generated content received 52.3, an average deviation of -4.9% from true submissions.
- Summary results: Detected papers averaged 51.7% compared with 52.6% for undetected papers, with no significant difference in grades.
- Summary results: Faculty feedback ranged from praise for well-supported ideas and clear thinking to criticism of insufficient depth, weak focus, confusing style, and poor referencing.
- Summary results: Markers also noted missing personality or visuals, lengthy introductions, source problems, and failure to address learning outcomes.
Discussion
The study finds that Turnitin can flag many GPT-4 submissions, but adversarial prompting limits the amount of AI-generated content detected and faculty do not consistently act on the results. The findings support assessment redesign, faculty training, and continued improvement of detection software.
- Detection and academic integrity: 91% of GPT-4 submissions were identified as containing some AI-generated content, but Turnitin detected only 54.8% of the total AI-generated content.Adversarial prompt engineering was used to evade detection.
- Assessment outcomes: Similar average scores between detected and undetected papers suggest that detecting AI-generated content did not significantly influence grading.Assessment requirements and design were associated with lower grades in some cases.
- Detection and academic integrity: Faculty formally reported 54.5% of experimental papers as potential academic misconduct cases, despite Turnitin identifying AI content in 91% of submissions.Some papers with high AI scores were not reported.
- Faculty practice: Faculty training is needed to interpret Turnitin results alongside indicators such as overly complex language, missing course content, and falsified references.The authors argue that detection tools are not infallible and should be combined with training.
- Faculty practice: Three markers correctly identified test submissions, while smaller cohorts may help instructors recognize students’ writing capabilities and progress.The study also identifies group submissions as a potential way to improve detection.
- Assessment outcomes: A paper with an 80% Turnitin AI score received 75 out of 100, indicating that AI-generated work may still be judged valuable or well-structured.This raises concerns about assessment reliability when detection results are not acted upon.
- Assessment outcomes: Assessment tasks requiring specified frameworks, pre-approved topics, or datasets produced very low scores in several cases, even without detected AI use.These requirements may create barriers to successful AI-generated submissions.
- Implications: The authors recommend assessment strategies that discourage undetectable AI use or deliberately incorporate AI, rather than relying solely on an unsustainable detector–generator arms race.They also call for ongoing collaboration between academia and detection-software developers.
Conclusion
The study provides initial insights into detecting AI-generated content in university assessment, while highlighting limitations in detection accuracy, accessibility, and generalizability. It recommends continued research, improved assessment strategies, and training to balance responsible AI use with academic integrity.
- Conclusion: Turnitin showed potential to support detection, but participants’ relatively low detection accuracy underscored the need for further training and awareness.The study also reports no significant difference between mean scores for AI-generated and student submissions.
- Conclusion: Unequal access to detection tools and AI technology raises concerns about their viability and accessibility across higher education contexts.The study identifies socioeconomic disparities in digital access and AI exposure as an area requiring further attention.
- Conclusion: Assessment tasks that explicitly involve AI tools, together with comprehensive faculty and student training, could promote responsible use while preserving assessment integrity.The paper presents this as a future direction for higher education assessment.
- Conclusion: Further research is encouraged to develop robust strategies that use AI’s potential while preserving academic integrity.The authors identify this as an ongoing need as AI continues to evolve.
Statements and Declarations
The authors disclose no study funding, relevant institutional employment connections, and ethics approval with informed consent. They also explain that full marking commentaries cannot be shared and that GPT-4 supported manuscript preparation.
- Statements and Declarations: No funding was received for conducting the study.
- Statements and Declarations: Several authors were employed by, or previously employed by, the University where the study took place.
- Statements and Declarations: The study received Human Ethics Committee approval, and participants gave informed consent with options to opt out at any time.
- Statements and Declarations: Full academic staff commentaries cannot be shared because of assessment-result confidentiality, while other study data are included in the article.
- Statements and Declarations: GPT-4 accessed through ChatGPT supported manuscript preparation, with the authors retaining responsibility for the content and views.