Source-linked AI summary

Could an Artificial-Intelligence agent pass an introductory physics course?

Gerd Kortemeyer

arXiv:2301.12127v2physics.ed-ph

TL;DR

The paper asks how human-like ChatGPT’s physics reasoning is and whether it could pass an introductory physics course. It evaluates the system on representative assessments from an actual calculus-based course, grading responses as human work would be graded. ChatGPT would receive a 1.5 grade, enough for course credit, while exhibiting many beginning-learner preconceptions and errors.

  • Problem

    The study asks whether ChatGPT’s human-like language extends to the logical, conceptual, mathematical, and strategic demands of introductory physics.

  • Method

    The authors test ChatGPT on representative assessments from an actual calculus-based physics course and grade its responses using human-like course criteria.

  • Results

    ChatGPT would receive a 1.5 grade, sufficient for course credit but below the grade-point average required for a bachelor’s degree.

  • Takeaways & Limitations

    ChatGPT’s behavior resembles an articulate but unstable beginning physics learner, including convincing presentation of incorrect physics and failures in evaluation strategies.

Abstract

from arXiv · show

Massive pre-trained language models have garnered attention and controversy due to their ability to generate human-like responses: attention due to their frequent indistinguishability from human-generated phraseology and narratives, and controversy due to the fact that their convincingly presented arguments and facts are frequently simply false. Just how human-like are these responses when it comes to dialogues about physics, in particular about the standard content of introductory physics courses? This study explores that question by having ChatGTP, the pre-eminent language model in 2023, work through representative assessment content of an actual calculus-based physics course and grading the responses in the same way human responses would be graded. As it turns out, ChatGPT would narrowly pass this course while exhibiting many of the preconceptions and errors of a beginning learner.

I. INTRODUCTION

The introduction asks whether ChatGPT can handle the logical, conceptual, mathematical, and strategic demands of introductory physics, testing it against representative course assessments. Its dialogue resembles an articulate beginning learner, with omissions, plug-and-chug calculation, missing units, and accumulating rounding errors.

  • Its primary physics impact is framed as a test of logical, conceptual, mathematical, and strategic problem-solving abilities rather than cheating.
  • The sample dialogue resembles an office-hour exchange between an instructor and a beginning physics student.
  • ChatGPT initially fails to consider that a car may have changed direction when asked how far it is from its starting point.When prompted, it recognizes that information is missing.
  • ChatGPT uses plug-and-chug calculations, omits units, and accumulates rounding errors instead of recognizing when speed cancels.These calculation weaknesses are characterized as shared with beginning physics learners.
  • ChatGPT is evaluated on representative assessment components from an introductory calculus-based physics course and compared subjectively with human learners.
  • ChatGPT cannot learn permanently by attending the course because it is pretrained, although OpenAI continues training it using user interactions.Learning-like behavior can occur within individual dialogues without producing permanent learning beyond them.

II. SETTING

The study examines first-year calculus-based physics lecture courses covering standard mechanics, thermodynamics, electricity and magnetism, and introductory modern physics. Laboratories are separate courses in the sequence, and course materials were drawn from multiple years.

  • The setting is first-year calculus-based physics lecture courses previously taught by the author at Michigan State University.Materials were gathered from different years of the same course to enable comparison with prior studies.
  • The first semester covers mechanics, including rotational dynamics, and the beginnings of thermodynamics.
  • The second semester covers electricity and magnetism plus introductory quantum physics and special relativity.
  • First- and second-semester laboratories are separate courses in the sequence.

III. METHODOLOGY

The study empirically scores ChatGPT across varied assessment formats using course-like conditions, with different rules for dialogue, attempts, grading, and input representation. Its methodology is described as strictly empirical and arguably anecdotal.

  • ChatGPT’s performance is assessed across different assessment problems, with scoring adapted to each component’s function in the course.
  • The Force Concept Inventory is scored by agreement with the answer choice.
  • Homework permits multiple attempts and dialogue, while clicker questions replay an actual lesson and allow peer-instruction discussions.
  • Programming exercises use course grading criteria with dialogue allowed, whereas exams prohibit dialogue and count only the first answer.
  • Because ChatGPT responses are probabilistic, the study generally evaluates the first dialogue and restarts after system errors or incorrect prompts.System errors occurred in about one-in-ten dialogues and were treated as platform overload rather than dialogue-specific failures.
  • Original figures and graphs are transcribed into text, substantially changing the character of those problems.This alteration is described as unavoidable for the text-based tool.
  • The methodology is strictly empirical and arguably anecdotal, although the course is described as typical of introductory physics courses worldwide.

IV. RESULTS

ChatGPT reaches the suggested entry threshold for Newtonian physics on the concept inventory, while displaying common novice misconceptions and logical errors. Its errors include impetus reasoning, confusion between individual and net forces, unstable concepts, and incorrect final conclusions.

  • ChatGPT scores 18 out of 30 points, or 60%, on the Force Concept Inventory.This corresponds to the suggested entry threshold for Newtonian physics and performance of a beginning learner who has grasped basic classical-mechanics concepts.
  • A modified final inventory question indicates that changed scenarios and answer order do not determine ChatGPT’s responses.The passage uses this result to distinguish ChatGPT from a novice relying on surface features.
  • ChatGPT exhibits the impetus preconception that an object moves in the direction of an applied force regardless of its initial movement.It also sometimes assumes that removing the force restores the original motion.
  • ChatGPT confuses individual forces with net force, shows unstable concepts, and sometimes follows a correct strategy but reaches an incorrect final conclusion.

B. Homework

ChatGPT solved 55% of assigned homework problems, but numerical errors and persistent formula-manipulation difficulties limited its performance. Its tendency to continue guessing after mistakes resembled beginning learners’ lack of reflection.

  • Numerical errors: Calculation errors occurred on 25 of 51 numerical problems, and ChatGPT usually failed to recover after errors were identified.The paper relates this pattern to language-model processing by pattern matching rather than directly manipulating equations.
  • Numerical errors: Adding “explain each step separately and clearly” sometimes prompted step-by-step calculation with intermediate results and helped overcome numerical problems.The passage describes this as anecdotal evidence rather than a systematic intervention.
  • Performance: ChatGPT solved 55% of 76 homework problems using an average of 1.88 attempts.The homework covered trajectory motion, friction, thermodynamics, capacitance, and special relativity; one diagram-dependent relativity problem was omitted.
  • Performance: Performance varied by problem set: 48% for trajectory motion and friction, 68% for thermodynamics, 62% for capacitance, and 36% for special relativity.The paper attributes the differences mainly to mathematical difficulties, especially manipulating formulas involving square roots.
  • Error recovery: After making a mistake, ChatGPT was unlikely to recover and often repeated the same or random errors across subsequent attempts.This pattern resembles students who repeatedly try the same approach without reflecting on what went wrong.

C. Clicker Questions

ChatGPT correctly solved 10 of 12 clicker questions, corresponding to a 93% course score under the course’s participation-credit scheme. It maintained some correct answers despite intentionally confusing discussion, but also made late-stage algebraic errors.

  • Individual questions: ChatGPT solved X1, X10, X11, and X12 correctly, although X10 required correcting a false start during the derivation.The resulting response resembled a stream-of-consciousness monologue.
  • Peer instruction: After peer instruction, ChatGPT answered all three repeated questions correctly and maintained its original answer despite an intentionally confusing discussion.The repeated questions were X5, X6, and X7, corresponding to X2, X3, and X4.
  • Calculation errors: For X8 and X9, ChatGPT set up the equations correctly but made a final sign error in X8 and dropped a factor 2 in X9.Both mistakes led to inconsistent or incorrect answer choices.
  • Overall performance: ChatGPT correctly solved 10 out of 12 clicker questions, yielding a 93% score under the course’s 60%/100% participation-credit scheme.The score was higher than most students achieved, although students were still learning the concepts while ChatGPT was not learning during the assessment.

D. Programming Exercises

ChatGPT completed the programming exercise with instructor-style feedback and earned full credit plus the available bonus, for 120%. Its initial code contained conceptual errors that one corrective comment fixed.

  • Code development: ChatGPT initially added the initial velocity at every time step and used the Coulomb force in the opposite direction.A single user comment corrected both errors, analogous to feedback from instructors or fellow students.
  • Bonus task: ChatGPT added the x-position graph after a third prompt, although the simulation itself could not run inside ChatGPT.The code could be copied into an environment such as a Jupyter Notebook for execution.
  • Performance: ChatGPT earned 120% in the programming component, receiving full credit and the 20% bonus.The result exceeded the performance of students in the course despite their extensive collaboration opportunities.

E. Exams

ChatGPT scored 47% on the mechanics final exam, with grading analysis indicating that numerical errors and flawed reasoning affected several answers. Repeated prompts could also produce different correctness outcomes, underscoring the probabilistic nature of its responses.

  • Exam performance: 47%: ChatGPT answered 14 of 30 mechanics final-exam questions correctly.The exam used answer options for human students, but ChatGPT received no options.
  • Grading analysis: Five answers were wrong because of numerical-calculation errors, while five correct answers relied on flawed reasoning and would not have received full credit.Hand grading would therefore distinguish answer correctness from the quality of the solution process.
  • Response variability: The same thermodynamics problem was solved correctly on the exam but incorrectly as homework, despite the homework allowing multiple attempts and help.The altered numerical values did not eliminate the broader inconsistency between responses to related problems.
  • Course implication: Exam-only grading would have produced a 1.0 course grade, barely enough for credit and below the 2.0 GPA required for graduation.This estimate uses the course’s 0.0-to-4.0 grading scale.

F. Course Grade

Using representative course components, ChatGPT’s weighted score was 54.55%, enough for course credit but below the graduation-grade threshold. Better numerical operations would have raised the estimate to 60%.

  • Clicker component: Three clicker items were presented before and after peer discussion in the replayed lecture assessment.The repeated items were part of the clicker component used in the course-grade estimate.
  • Weighted course grade: 54.55%: weighting homework, clickers, programming, and exams yielded a course grade of 1.5, enough for course credit.The typical weighting was 20% homework, 5% clicker, 5% programming, and 70% exams.

V. DISCUSSION

The discussion portrays ChatGPT as articulate and often human-like, yet unstable in physics reasoning, numerical work, and self-evaluation. Its performance raises questions about assessment, computation-integrated curricula, and the human competencies needed for collaboration with AI.

  • Human-like behavior and limits: ChatGPT resembles an articulate but unstable undergraduate: it has rudimentary physics knowledge, calculator-like numerical weakness, and no metacognition.The discussion contrasts human-like dialogue with failures that reflect probabilistic language-model behavior.
  • Peer instruction: ChatGPT sometimes maintained a correct answer through intentionally confusing peer-instruction dialogue, so the repeated item remained counted as solved.The X3/X6 dialogue illustrates persistence rather than learning from the simulated discussion.
  • Confident errors: Its convincing presentation of incorrect physics may reflect training on physics text that was not uniformly correct.The paper connects human-like novice errors with the likely composition of the training corpus.
  • Programming performance: The programming exercise was an anomaly because ChatGPT’s language model clearly extended to programming languages.The authors contrast this success with the difficulties seen in other course work.
  • Computational curricula: Computation-integrated curricula intended to make physics problem solving more authentic may now face an on-demand program generator at an uncharted level.The discussion frames this as a challenge for emerging curricular efforts, not as a demonstrated causal effect.
  • Implications for education: A course passed by a language model can prompt educators to reconsider what conceptual understanding students need for working with AI.The discussion presents this as a wake-up call rather than a settled redesign prescription.
  • Possible responses: Educators may respond with detection tools, high-stakes proctored exams, or a broader reconsideration of course goals.The discussion notes tensions between high-stakes testing and research favoring frequent formative assessment and spaced repetition.
  • Human competencies: Physics education may need to emphasize evaluating one’s own work, including dimensional analysis, estimates, coherence checks, implications, and limiting cases.The paper identifies self-evaluation as a human skill that AI is very unlikely to perform reliably.

VI. CONCLUSION

ChatGPT would have earned course credit with a 1.5 grade in the studied introductory physics sequence, while improved numerical algorithms would have raised the estimate to 2.0. Its lack of metacognition means it presents truthful and misleading information with equal confidence, shifting attention toward the human skills needed for collaboration with AI.

  • Conclusion: ChatGPT would have earned a 1.5 grade in the standard introductory physics lecture sequence, enough for credit but below the graduation GPA requirement.The study’s counterfactual grade is based on the course’s representative assessment components.
  • Conclusion: Better algorithms for simple numerical operations would have raised the estimate to 2.0, enough for graduation if performance generalized across other courses.This is explicitly conditional on similar performance elsewhere.
  • Conclusion: Because ChatGPT lacks metacognition, it presents truth and misleading information with equal confidence.The paper therefore frames the central educational challenge as identifying inherently human skills and competencies for future AI collaboration.
Loading 2301.12127v2…