Source-linked AI summary
The Death of the Short-Form Physics Essay in the Coming AI Revolution
Will Yeadon, Oto-Obong Inyang, Arin Mizouri, Alex Peach, Craig Testrow
TL;DR
This paper asks whether readily accessible AI systems can undermine short-form Physics essays as an assessment method. It generates and evaluates ten five-essay submissions using an accredited university module’s assessment framework. The submissions averaged 71 ± 2%, matching the module average of 71 ± 5%, while plagiarism scores remained low; the authors therefore identify a serious assessment-fidelity threat, within the study’s stated scope and comparison limits.
Problem
AI systems can produce accurate, human-like text cheaply and quickly, while universities generally check plagiarism rather than whether submitted work was AI-generated.
Method
The study generated ten five-question AI submissions using varied prompts and evaluated them against an accredited Physics module’s assessment framework.
Results
71 ± 2% was the average mark for ten AI-generated submissions, compared with 71 ± 5% for the Physics in Society module; plagiarism scores were 2 ± 1% on Grammarly and 7 ± 2% on TurnitIn.
Takeaways & Limitations
The results indicate that AI-generated short-form Physics essays can attain marks comparable to second-year Physics students and challenge the fidelity of this assessment method.
Takeaways & Limitations
The study compares AI scores with module averages rather than student submissions for the same exam, and focuses on short-form essays rather than broader STEM tasks.
Abstract
from arXiv · showhide
The latest AI language modules can produce original, high quality full short-form ($300$-word) Physics essays within seconds. These technologies such as ChatGPT and davinci-003 are freely available to anyone with an internet connection. In this work, we present evidence of AI generated short-form essays achieving first-class grades on an essay writing assessment from an accredited, current university Physics module. The assessment requires students answer five open-ended questions with a short, $300$-word essay each. Fifty AI answers were generated to create ten submissions that were independently marked by five separate markers. The AI generated submissions achieved an average mark of $71 \pm 2 \%$, in strong agreement with the current module average of $71 \pm 5 %$. A typical AI submission would therefore most-likely be awarded a First Class, the highest classification available at UK universities. Plagiarism detection software returned a plagiarism score between $2 \pm 1$% (Grammarly) and $7 \pm 2$% (TurnitIn). We argue that these results indicate that current AI MLPs represent a significant threat to the fidelity of short-form essays as an assessment method in Physics courses.
1. Introduction
AI text-completion systems can produce accurate, critical, human-like responses cheaply and quickly, creating a potential threat to conventional assessment. This study examines whether such systems can generate university-level short-form Physics essays.
- Background: AI text-completion technologies can reliably produce accurate, clear and critical content on practically any topic from a brief prompt.The systems are described as increasingly accessible, fast, cheap and easy to use.
- Assessment threat: AI-generated text may pass plagiarism checks and receive higher marks than a student’s unaided work.Universities currently look for plagiarism rather than whether text was generated by AI.
- Study aim: The study argues that AI-generated short-form Physics essays can achieve First Class marks in an accredited university module.This claim motivates testing AI essays against an assessed Physics assignment.
- Prior work: Newer systems such as ChatGPT and davinci-003 can generate entire essays from a single user-defined sentence prompt.Earlier GPT-2 work involved blending student writing with model output, whereas newer systems can produce complete essays.
- Model capabilities: davinci-003 output can demonstrate critical understanding and reasoning in responses to essay questions.The paper illustrates this with literary analysis and an apparent moral position on using AI to generate essays.
2. Method
The study uses a five-question, short-form Physics assessment and generates unedited AI answers through varied prompting before evaluating their content against the module’s marking framework.
- Assessment: The Physics in Society exam contains five short-form essay questions of no more than 300 words covering history, philosophy, communication and ethics of Physics.The questions were used to generate the AI submissions.
- Assessment: The assessment proforma evaluates essays using five key criteria.The criteria are represented in the Physics in Society grading proforma.
- Assessment: Students receive equally weighted marks across five assessment categories, with each category judged from answers to all five questions.The module’s typical score is 71 ± 5.
- Generation: A sample of n = 10 AI-generated scripts was compiled from davinci-003 outputs based on the assessment questions.Each script contained five question-answer pairs.
- Generation: AI responses were generated by rephrasing questions and adding prompts requesting discursive essays of specified lengths.Prompt variation was used to avoid brief, laconic responses and produce more original answers.
- Generation: The AI output was not edited, except that excessively similar generations were rejected and replaced.This provided a consistent benchmark of the model’s generated essays.
3. Analysis and results
AI-generated essays matched the module average and were difficult to distinguish from student work through marking and plagiarism checks, although their raw text contained minor readability and language imperfections.
- 71 ± 2% was the average mark for ten AI-generated submissions, matching the Physics in Society average of 71 ± 5%.
- Most students performed comparably to or worse than the AI essays, while the very highest-performing students outscored them.
- 2 ± 1% plagiarism on Grammarly and 7 ± 2% on TurnitIn would both be deemed sufficiently original for university assessment.
- Independent marker scores overlapped, with marker averages ranging from 69 ± 2% to 73.0 ± 1.6%.The overall independent-marker average was 71 ± 2%.
- The AI output included phrases whose readability could be improved, showing that raw generated text was not always semantically perfect.
- American English spelling required changes for a UK context, but the authors state that this imperfection would not prevent use as an essay-writing tool.
4. Discussion & Conclusion
The study finds that AI can produce high-quality Physics essays and argues that this challenges non-invigilated short-form essay assessment. The authors discuss practical responses, limitations of the evidence, and possible future applications of AI in education.
- Invigilated settings are presented as a simple practical way to reduce misuse without completely redesigning existing assessments.
- The study uses a rudimentary comparison with the module average rather than student submissions, and future work will compare AI and human work and assess marker discrimination.
- The present work covers only short-form essay questions, while future work will examine reports and analytical tasks involving calculations, coding, symbolic manipulation, and algebraic typesetting.
- AI-generated feedback can be specific and rubric-related, although ChatGPT may score generously compared with the rubric.
- The authors argue that non-invigilated short-form essay assessments are vulnerable to current AI text-completion technologies.