Source-linked AI summary

Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes

Praveen Pathak, Siddharth Tiwary, Charudatt Kadolkar, Vijay Singh, David Rakestraw, Shirish Pathare, Anwesh Mazumdar

arXiv:2608.20521v1physics.ed-phcs.AIcs.CY

TL;DR

High-stakes handwritten physics grading requires AI scores to agree with official marks and assessment outcomes. This study evaluates rubric-based multimodal AI grading across physics examinations and finds strong score agreement and recovery of the official Olympiad team, while exact partial-credit grading remains difficult.

  • Problem

    It remains unclear whether AI grading agrees closely enough with official scores to support high-stakes handwritten physics assessment outcomes.

  • Method

    The study retrospectively compares two rounds of rubric-based AI grading with official grading across handwritten physics examinations and selection assessments.

  • Results

    AI showed strong agreement with official total scores and recovered the same top-five Olympiad selection group as human examiners.

  • Takeaways & Limitations

    AI is best used as a second reader or targeted audit tool, with expert examiners retaining responsibility for rubrics and borderline cases.

  • Takeaways & Limitations

    The study used moderated official scores as the human reference and did not conduct a separate human–human regrading experiment.

Abstract

from arXiv · show

Multimodal AI can read handwritten physics solutions, but high-stakes grading requires agreement with official scores and outcomes. This study evaluated GPT-5.5-based grading on 10364 scanned pages from 520 handwritten submissions by 416 unique candidates or students across three assessments: a national Physics Olympiad theory examination, the final Olympiad selection camp with theory and experiment components, and a university quantum-mechanics examination. Each submission was graded twice by AI using the official rubrics. The second round used revised page-by-page and evidence-location instructions developed after first-round disagreement analysis. During grading, AI did not see official human marks or AI--human comparisons. Total-score correlations with official marks were high (0.91--0.97). For the final Olympiad selection, AI recovered the same five-student team as official grading. The second round improved aggregate question-part agreement, especially where first-round disagreements were larger. The main difficulty remained exact partial-credit grading, especially in experimental work. Reliable AI grading therefore depends on detailed rubrics and should be used as a second reader or audit tool under examiner control.

I. INTRODUCTION · II. DATA AND METHODS · A. Assessments and reference scores

The study evaluates whether vision-capable AI can grade handwritten physics examinations closely enough to support high-stakes assessment decisions. It spans Olympiad selection and university quantum-mechanics assessments, using official examiner scores as reference outcomes and examining how grading procedures affect agreement.

  • I. INTRODUCTION: Small score differences matter because Olympiad grading can change candidate rankings, team selection, and medal awards.The questions are long and often require judgment beyond matching a textbook answer.
  • I. INTRODUCTION: The assessment set covers a national Olympiad theory examination, a final theory-and-experiment selection camp, and a university quantum-mechanics examination.Together, these tasks include theory, experiment, derivation, data analysis, diagrams, and conceptual reasoning.
  • II. DATA AND METHODS: The analysis compares AI and official scores at total-score, selection, question-part, and question-type levels, while examining representative disagreements and experimental grading.AI grading was retrospective, after official selection results and university grades had been released, so it did not alter those outcomes.
  • A. Assessments and reference scores: All submissions were handwritten, and the study processed 10 364 scanned pages from Olympiad examinations conducted at two stages and 40 university quantum-mechanics submissions.The Olympiad datasets came from different years because of submission availability and scanning history.
  • A. Assessments and reference scores: Official reference scores were final marks checked by a second examiner, with concerns reconsidered before release and student review mechanisms available afterward.OE1 permitted regrading requests, while OE2 and QM allowed students to inspect graded work and discuss awarded points with examiners.
  • A. Assessments and reference scores: OE2’s official ranking combined theory and experiment using a 60:40 weighting, equivalent in normalized analyses to 0.6T + 0.4E.The combined score was constructed from prorated theory and experiment components totaling 400 marks.

B. AI grading procedure and grading rounds

AI graded anonymized handwritten submissions with the official questions, solutions, rubrics, and instructions, without access to human scores or outcomes. A second full grading round revised instructions based on first-round disagreement analysis, while focused tests examined whether clearer physics and scoring conditions reduced discrepancies.

  • AI grading procedure: The AI received exam questions, solutions, official rubrics, and anonymized handwritten submissions after visible human scores and comments were erased.The same official rubrics were used for human and AI grading, although human examiners interpreted them case by case.
  • Round I: Round I was an initial full AI grading run with outputs containing scores and comments for every question part or subpart, without access to human scores, totals, ranks, or selection status.The run used the submissions, questions, solutions, rubrics, and general grading instructions.
  • Focused refinement: Focused refinement tested three repeated or large disagreements—one each from QM, OE1, and OE2—to assess whether clearer physics and scoring conditions reduced AI–official score differences.These diagnostic runs informed the revised instructions used in Round II.
  • Round II: Round II was a new full grading run that revised instructions after Round I disagreement analysis while withholding all prior scores, comments, ranks, selection status, and AI–human comparisons.The remaining parts retained Round I rubrics, while three repeatedly or substantially disagreeing parts used refined rubrics.

C. Evaluation measures · III. AGREEMENT AND SELECTION RESULTS · A. Total-score agreement

AI and official total scores agreed strongly across all three assessments, with Pearson correlations of 0.91–0.97 in RI and 0.93–0.96 in RII. The revised workflow reduced positive AI over-scoring and absolute score differences, especially for OE1 and QM.

  • C. Evaluation measures: The study made 7058 AI–official comparisons, each pairing one student response with one official question part or subpart across OE1, OE2, and QM.Each part or subpart had its own maximum mark and specified how marks were awarded or deducted.
  • C. Evaluation measures: D and MAD were reported as percentages of the relevant maximum possible score unless raw marks were explicitly stated.Table III defines the analysis quantities used for the agreement results.
  • A. Total-score agreement: r = 0.91–0.97 in RI and r = 0.93–0.96 in RII, with scores tracking closely across all three assessments.Fig. 1 compares OE1, QM, and OE2 theory and experiment totals; points lie close to the equal-score diagonal in both rounds.
  • A. Total-score agreement: RII was a full revised workflow combining page-by-page checking, evidence notes, stricter permitted-score checks, confidence and review fields, and clearer selected-question instructions.The three focused refinements later discussed represented only about 3.6% of the combined official rubrics, so RI–RII differences reflect the full workflow.
  • A. Total-score agreement: OE1 MAD fell from 7.1% to 4.8%, while QM MAD fell from 9.4% to 3.8% after the revised workflow.The revised workflow also reduced the positive AI–human shift in OE1 and QM.
  • A. Total-score agreement: RI produced positive D in all four total-score comparisons, indicating that AI generally awarded more points than human examiners.RII reduced this upward shift in OE1 and QM.
  • A. Total-score agreement: The residual plots show that RI differences were more often above zero, whereas RII reduced the upward shift across the score range.The differences were not confined to low-scoring submissions or scores near a cutoff.

B. Top-cohort and selection agreement

AI recovered the official five-student OE2 team in both grading rounds and matched broader selection groups closely. Agreement also improved for OE1 top-cohort overlap and QM course grades, though exact ranks near cutoffs still required human grading.

  • OE1 top-cohort agreement: For OE1, each round recovered 31 of the human top 40 and 40 of the human top 50, while top-10 overlap increased from 3/10 in RI to 7/10 in RII.This agreement supports identifying the top group, but human grading remains necessary for exact ranks near a cutoff.
  • OE2 selection agreement: The AI top five contained the same five students as the human top five in both OE2 rounds, while ranking them in a different order.At k = 10, both rounds recovered 9 of the human top 10; at k = 20, both recovered 19 of the human top 20.
  • QM course-grade agreement: RII exactly matched 34 of 40 official QM course grades, compared with 26 of 40 in RI, and all 40 RII grades were within one grade step.The released grades span nine categories, from FP to AS, with AS the highest grade; these comparisons are shown in Fig. 3.
  • OE2 mode dependence: The Pro Standard OE2 run recovered the same human top-five group as the main Thinking High run, although the internal order changed.At broader cutoffs, Pro Standard recovered 9 of the human top 10 and 19 of the human top 20.

C. Question-part agreement and partial credit · IV. STRENGTHS AND DISAGREEMENTS IN AI GRADING

RII improved question-part agreement over RI, raising exact matches and reducing large errors across 7058 parts. The revised grading also improved broad credit-band recognition and partial-credit calibration, though exact partial-credit amounts remained difficult.

  • C. Question-part agreement and partial credit: RII raised exact agreement from about 63% to about 70% across 7058 official question parts and cut differences larger than one point from about 13% to 7%.Question parts were compared using raw-point differences d = |AI −Human|, with bands based on the usual 0.5-point scoring increment.
  • IV. STRENGTHS AND DISAGREEMENTS IN AI GRADING: At the experimental-component level, RII achieved 77% exact agreement, with 16% differing by at most 0.5, 4% by 0.5–1, and 3% by more than 1 raw mark.This finer-grained check was for analysis only because some official experiment parts combined several judgments into one mark.
  • C. Question-part agreement and partial credit: RII placed 80% of human-zero, 86% of human-partial, and 87% of human-full parts in the corresponding broad credit bands.Partial-credit parts represented 40% of available points, making band-level recognition especially consequential.
  • C. Question-part agreement and partial credit: Exact agreement on human-partial parts increased from 24.6% in RI to 32.8% in RII, while conditional MAD decreased from 0.92 to 0.72 raw marks.These results show improved calibration when both the human examiner and AI assigned partial credit.
  • IV. STRENGTHS AND DISAGREEMENTS IN AI GRADING: When AI moved a human-zero part into the partial-credit range, its average award was about one raw mark, and this occurred less often in RII than RI.Thus, the revision improved both broad category recognition and partial-credit calibration.
  • IV. STRENGTHS AND DISAGREEMENTS IN AI GRADING: The analysis proceeds from aggregate question-part results to representative successes, experimental grading, question-type differences, and recurring disagreement patterns.These analyses contextualize the observed agreement and partial-credit results.

A. Examples of physics-specific grading · B. Experimental grading

The AI often performed physics-specific grading by locating relevant work across pages, interpreting derivations and sketches, and applying rubric-based partial credit. Experimental grading was strongest when evidence was explicit, but human review remained important for judging consistency among data, fits, uncertainties, and conclusions.

  • A. Examples of physics-specific grading: The AI often located relevant solutions across multipage submissions, interpreted handwritten derivations, credited equivalent mathematical forms, and identified specific physics errors.These physics-specific comments supported the observed total-score agreement.
  • A. Examples of physics-specific grading: The AI credited correct detailed working despite a factor-of-two transcription error in the summary answer box.The detailed working contained the correct self-energy factor and final result, illustrating why full-submission inspection matters in Olympiad-style grading.
  • A. Examples of physics-specific grading: The AI found required effective-potential sketches outside the designated answer box and awarded full credit when their qualitative features matched the rubric.It also distinguished limited credit when a sketch showed only some required endpoint behavior.
  • B. Experimental grading: Experimental grading required reading tables, calculations, graphs, fit lines, uncertainty estimates, and written conclusions, with official components sometimes corresponding to more granular judgments.RII recorded dimensions such as table quality, graphing, fit or slope extraction, uncertainty, and conclusion separately.
  • B. Experimental grading: The AI often recognized transformed tables, plotted points, fit and limiting lines, slopes, and reported values.Its weaker cases involved whether the data range, fit, uncertainty, and final conclusion supported one another.
  • B. Experimental grading: The AI identified fit and limiting lines as grading evidence and agreed with partial credit when plotted points alone did not satisfy the full graph-analysis requirement.The examples included agreement on data-table and graph-analysis components as well as a partial graph lacking usable fit, slope, or uncertainty construction.
  • B. Experimental grading: Human review remained important when experimental marks depended on physical consistency among the table, graph, fit, uncertainty, and conclusion.The examples indicate that multimodal AI can often read and score experimental evidence, but cross-component consistency is a persistent review condition.

C. Agreement across question types

Agreement between AI and official scores extended across numerical, derivation, conceptual, diagram, and data/graph question types. Revised second-round grading reduced mean absolute differences for every type, especially when credit depended on evidence satisfying a specific rubric criterion.

  • Agreement across question types: Question-type labels grouped official parts by their primary grading judgment, with OE2 experimental sections labeled by their main official task.The labels were broad but enabled consistent comparison across OE1, OE2, and QM.
  • Agreement across question types: AI scores generally increased with human scores across derivations, numerical answers, conceptual reasoning, diagrams, and data/graph tasks.Conceptual written reasoning showed the greatest over-scoring in RI, while figure-data and figure-diagram tasks depended on satisfying specific visual or experimental evidence requirements.
  • Agreement across question types: Reviewed cases suggest that page-by-page checking and clearer credit conditions improve grading most when required evidence appears in a specific response part.Across formats, the central challenge was determining whether a table, graph, diagram, or derivation met the rubric’s specific physical criterion.
  • Agreement across question types: MAD was lower in RII than RI for every question type across the 7058 official question parts.The comparison used MAD as a percentage of each part’s maximum possible score.
  • Agreement across question types: RII improvements extended beyond explicitly refined questions, including figure-data, figure-diagram, numerical, conceptual, and derivation parts.The largest gains appeared in figure-data and figure-diagram parts, but improvement also occurred in numerical, conceptual, and derivation tasks.

D. Recurring sources of disagreement · V. IMPROVING AND REVIEWING AI GRADES · A. Focused rubric refinement

Review of recurring disagreements motivated clearer scoring conditions and human-review flags, while focused rubric refinements reduced disagreement by making required physics and credit conditions explicit. These results support detailed, physics-specific rubrics for reliable AI grading.

  • D. Recurring sources of disagreement: Recurring disagreement patterns led to two proposed safeguards: clearer scoring conditions and confidence flags for human review.The larger disagreements had identifiable causes summarized in Table VII, motivating focused refinement and later revised full grading.
  • A. Focused rubric refinement: Focused refinements reduced several AI–official disagreements by explicitly stating the required physics and scoring conditions in QM 3(b), OE1 Q2, and OE2 T2 Q1(f).The cases were selected after reviewing Round I disagreements and regraded with question-specific instructions.
  • A. Focused rubric refinement: For the thermodynamic-cycle example, refined instructions stopped AI from awarding shape credit when the physical curvature was incorrect.The original rubric left intended curvature implicit, whereas the refined rubric stated the cycle’s physical conditions directly.
  • A. Focused rubric refinement: For the Grover-rotation diagram, RII matched the official 0.5 out of 3 points instead of RI’s 2 out of 3 by rejecting a plausible but nonrequired sketch.RII credited the qualitative oracle/diffusion idea but required the two-dimensional rotation diagram.
  • A. Focused rubric refinement: MAD fell from 1.1 to 0.4 marks, Pearson’s r rose from 0.70 to 0.82, exact agreement increased from 15% to 45%, and agreement within 0.5 mark rose from 35% to 80%.These results concern the Grover-rotation refinement example.
  • A. Focused rubric refinement: In the OE1 Q2 subset with large Round I differences, MAD fell from 38% to 17% of the part’s maximum score after requiring physical criteria and limiting unsupported conclusions.The refinement covered n = 72 submissions selected for large AI–official score differences.
  • V. IMPROVING AND REVIEWING AI GRADES: Reliable AI grading requires rubrics that explicitly define award conditions, common deductions, acceptable alternatives, and carried-forward-error rules before grading begins.The conclusion follows from the focused refinements showing that broad goals can leave physics needed for credit implicit.

B. Confidence and human-review flags

Confidence flags identified question parts with substantially larger grading deviations and can prioritize human review, but they did not reliably guarantee agreement with official scores. In OE2, some high-confidence parts still differed by more than one point, supporting review triage rather than automatic acceptance.

  • Confidence and human-review flags: MAD was much lower for high-confidence parts than for medium- or low-confidence parts across every reported comparison, including OE1 RII (6.0% vs 17.0%) and QM RII (9.8% vs 28.2%).In OE2, the corresponding theory comparison was 5.3% versus 27.1%.
  • Confidence and human-review flags: 87 of 1724 high-confidence OE2 question parts still differed from the official human score by more than one point.The OE2 analysis covered 1898 official question parts: 1482 theory and 416 experiment parts.
  • Confidence and human-review flags: Confidence flags should triage human inspection rather than authorize final acceptance without review.High-confidence parts help determine what humans inspect first, but some important disagreements remain undetected.

VI. DISCUSSION AND CONCLUSIONS

AI grading closely matched official outcomes across the examinations, including the same Olympiad top-five group and near-exact quantum-mechanics grades, but exact partial-credit decisions still require examiner oversight. Reliable deployment therefore depends on detailed rubrics, evidence-based page-level grading, and targeted human review.

  • Discussion and Conclusions: 34 of 40 quantum-mechanics grades were reproduced exactly, while all 40 were within one grade step; both AI rounds recovered the official Olympiad final top-five group.For OE2, the AI ranked the selected students differently despite recovering the same group; for OE1, it identified most of the larger top group.
  • Discussion and Conclusions: AI often recognized correct physics and relevant evidence, but remained more likely than human examiners to award partial credit for incomplete or visually plausible work.It could credit correct working despite summary-box transcription errors, accept equivalent forms, and identify important diagram features.
  • Discussion and Conclusions: Explicitly specifying physics conditions for credit improved grading in the QM Grover-diagram question and OE1 Q2, while OE2 changes were smaller and mixed.OE2 total-score agreement was already high in the first round, limiting the effect of refinement.
  • Discussion and Conclusions: High-confidence responses had smaller average score differences, but confidence fields still require examiner checking because some high-confidence responses disagreed with official markings.Confidence and review fields are therefore best used to prioritize manual review rather than to accept grades automatically.
  • Discussion and Conclusions: Page-by-page grading with evidence locations improved auditability and was associated with better human agreement for long submissions, although it increased grading time.Relevant work may appear outside expected answer spaces, making evidence localization particularly useful for lengthy handwritten responses.
  • Discussion and Conclusions: Expert examiners remain central: reliable use requires detailed rubrics, explicit credit and deduction conditions, and review of flagged or borderline cases.More intensive Pro modes are best suited to targeted audit, while AI shifts effort toward rubric definition and review; official scores were moderated and checked by multiple examiners.

DATA AND MATERIALS STATEMENT

Anonymized scores, task categories, prompts, and analysis scripts may be shared with approval, while raw handwritten submissions remain confidential.

  • Anonymized scores, task categories, prompts, and analysis scripts may be shared subject to approval, but raw handwritten submissions remain confidential because of privacy and examination constraints.
Loading 2608.20521v1…