Source-linked AI summary

Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education

Rayed AlGhamdi

arXiv:2609.05346v1cs.AI

TL;DR

The study addresses limited evidence on how students interpret AI-generated writing evaluation when its source is disclosed. Through qualitative analysis of student reflections, it finds that students value AI feedback but reserve grading authority for human instructors.

  • Problem

    Limited evidence exists on how students interpret AI-mediated writing evaluation when they know an AI system, rather than a human instructor, produced it.

  • Method

    The study used a course-embedded qualitative inquiry in which computing students reflected on ChatGPT’s disclosed evaluation of their handwritten writing.

  • Results

    Students found ChatGPT feedback technically useful and recognized its limitations, but distinguished feedback utility from evaluative authority and preferred human grading authority.

  • Takeaways & Limitations

    The findings support treating AI as a useful feedback mechanism while retaining the human instructor as the authority over grading decisions.

  • Takeaways & Limitations

    Findings are limited in transferability because participants were male computing undergraduates from one course section at one Saudi public university.

Abstract

from arXiv · show

The integration of GenAI tools into higher education assessment raises important questions about how students understand, interpret, and respond to AI-mediated evaluation. As instructors increasingly explore AI tools for providing feedback, prior research has examined whether GenAI-generated feedback improves writing performance and how students perceive its usefulness; comparatively little is known, however, about how students interpret such evaluation when they are explicitly informed that an AI system, rather than a human instructor, produced the feedback and the score. This study reports findings from a qualitative pedagogical inquiry conducted in an undergraduate technical communication course for computing students at a Saudi public university. Thirteen male undergraduate computing students completed an in-class handwritten writing task; the scanned submissions were evaluated by ChatGPT using a rubric-based prompt aligned with the task objectives. Students were then explicitly informed that ChatGPT had generated the score and feedback and were invited to reflect on the evaluation in writing. Inductive thematic analysis of these reflections identified four themes: perceived usefulness of feedback; awareness of AI's contextual and pedagogical limitations; conditional trust, distinguishing feedback utility from evaluative authority; and reflection on the institutional and pedagogical role of the human instructor. Participants accepted GenAI feedback as useful for surface-level revision but consistently positioned the human instructor as the appropriate authority over grading decisions. The study identifies this as a distinction between feedback utility and evaluative authority, two judgments that students treat as analytically separate rather than as opposite ends of a single approval scale...

1. Introduction

Writing is central to higher education and computing practice, while widespread student use of GenAI complicates assessment and creates opportunities for AI-assisted feedback. This study examines how students evaluate and reason about ChatGPT-mediated assessment when its evaluative role is transparent.

  • Student writing is important for coursework, employment, and computing practice, where graduates must explain technical concepts, document systems, and justify design choices.
  • Widespread student use of GenAI as a writing assistant creates uncertainty about authorship and developmental processes, making purely prohibitive responses impractical.
  • GenAI feedback may improve writing quality, increase engagement, and reduce instructor workload, particularly where large enrolments limit frequent individualised feedback.
  • The introduction identifies transparency as an underexamined issue because beliefs about whether feedback comes from an algorithm or human can shape perceived quality, fairness, and acceptance.
  • The study transparently informs students that ChatGPT evaluated handwritten writing, then elicits reflections on feedback, scores, writing insights, appropriateness, evaluator roles, and broader pedagogical concerns.The design complements earlier blinded inquiry, but direct causal comparison is not possible because cohorts, periods, and baseline AI familiarity differed.

2. Literature Review

Prior research shows that GenAI can provide technically adequate writing feedback, yet students interpret its usefulness, trustworthiness, and authority conditionally. This study addresses the gap by examining how Saudi computing students reason about AI-mediated evaluation when explicitly told that AI generated their feedback and score.

  • Technical adequacy: GenAI feedback can match instructor feedback in clarity and instructional usefulness, shifting research attention from technical adequacy toward students’ interpretations of AI-produced evaluation.Direct comparisons found similar student ratings when AI received a clear rubric, while systematic reviews frame GenAI as writing assistance and pedagogical support rather than instructor replacement.
  • Trust and evaluative authority: Students may appreciate AI for efficient, consistent, surface-level feedback while resisting its authority over high-stakes, subjective evaluation.Algorithm aversion and algorithm appreciation vary with error exposure, decision stakes, domain opacity, and perceived need for distinctly human judgment.
  • Feedback literacy and assessment: Feedback is a dialogical, interpretive process in which students judge, manage, and act on feedback while considering contextual constraints and the evaluator’s knowledge of their work.Feedback literacy frameworks emphasize student capacities and dispositions, and relational accounts highlight expectations about who evaluates, what they know, and what dialogue remains possible.
  • Reflective and metacognitive engagement: Students engage selectively with GenAI feedback, accepting, disputing, and contextualizing suggestions through cognitive and emotional tensions rather than receiving them passively.Prior work reports heterogeneous trust, skepticism, and conditional acceptance, motivating examination of whether explicit AI disclosure and summative grading intensify selective engagement.
  • Research gap and contribution: This study examines how Saudi computing students distinguish among AI’s assessment functions after being explicitly informed that AI generated their writing feedback and score.The analysis focuses on categories of reasoning emerging from transparent, post-assessment reflections rather than predefined acceptance constructs or a single accept-versus-reject judgment.

3. Methodology

This course-embedded qualitative inquiry examined students’ responses to transparent ChatGPT evaluation through handwritten writing and reflective responses analyzed inductively. The design used authentic in-class writing, rubric-based AI scoring, explicit disclosure, and thematic analysis of students’ reflections.

  • Research design: The study used a course-embedded pedagogical inquiry to examine how students respond when informed that ChatGPT, rather than the instructor, evaluates their work.The inquiry addressed AI-mediated assessment’s implications for students’ metacognitive awareness, trust, skepticism, ethical reasoning, and engagement with feedback.
  • Participants and setting: The sample comprised 19 male second-year computing students in one Saudi public-university Technical Communication course, with 13 reflections returned for analysis.The course emphasized clarity, logical organization, and audience awareness, and the reflective activity was embedded in routine assessment.
  • Study procedure: Transparency was the central intervention: students were explicitly told that ChatGPT had generated their feedback and score before completing a reflective writing task.The study procedure comprised writing, scanning, AI evaluation, disclosure, reflection, and inductive thematic analysis.
  • Writing task and AI evaluation: Students completed handwritten in-class writing, which was scanned and evaluated by ChatGPT using course-aligned criteria for clarity, organization, sentence correctness, and conciseness.Handwriting was deliberately used to ensure that the evaluated writing was generated by students rather than an AI system.
  • Data and analysis: The sole dataset consisted of spontaneous handwritten reflections addressing agreement with the AI evaluation, learning about writing, and opinions of ChatGPT as evaluator.Inductive thematic analysis followed Braun and Clarke’s guidelines through repeated familiarization and open semantic coding of student responses.

4. Findings

Among 13 students, ChatGPT feedback was valued for identifying specific writing problems, but participants questioned its contextual reliability and rejected it as the final grading authority. They instead positioned the human instructor as essential for dialogical, pedagogical, and institutional reasons, while noting the findings’ limited generalizability.

  • Perceived usefulness of feedback: Students consistently found ChatGPT feedback useful because it identified specific grammar, sentence-structure, and organization problems clearly rather than offering vague evaluation.Participants described the feedback as clear, direct, well-organized, and objective, and some reported increased confidence and learning.
  • AI’s contextual and pedagogical limitations: Participants questioned ChatGPT’s reliability when handwriting, writing intent, contextual meaning, institutional rating systems, or consistently positive evaluation could affect its judgments.These concerns led students to treat the evaluation as neither neutral nor infallible, especially in interpreting context and meaning.
  • Conditional trust: Students distinguished feedback utility from evaluative authority, accepting ChatGPT for learning support or initial feedback while rejecting it as the system that should decide grades.This conditional trust appeared even among enthusiastic participants and was treated as a separation between two judgments, not a single approval scale.
  • Human instructor’s role: Participants regarded the human instructor as essential because teachers know individual students, enable dialogue about circumstances, and safeguard the university’s pedagogical and institutional role in assessment.Human grading was framed as a relational practice involving interpretation, explanation, and the possibility of another chance.
  • Overall findings: The findings derive from a small, purposive sample of 13 students and therefore should not be generalized without further research.The study’s overall pattern was usefulness of AI feedback, critical awareness of limitations, preference for human grading authority, and recognition of the instructor’s safeguarding role.

5. Discussion

Students treated ChatGPT feedback as conditionally useful but not as authoritative grading, emphasizing that institutional and ethical responsibility should remain with human instructors. Transparency and post-assessment reflection broadened students’ reasoning from feedback accuracy to fairness, accountability, contextual judgment, and the role of AI in assessment.

  • Students distinguished agreement with ChatGPT’s useful feedback from acceptance of AI as an authoritative grader, positioning human instructors as responsible for grading decisions.Trust in AI outputs was contingent on instructor validation, interpretation, and oversight because students feared misinterpretation of context, intent, and effort.
  • The study’s contrast with an earlier blinded cohort suggests that disclosure may shape students’ boundaries around AI authority, although differing cohorts and contexts prevent direct comparison.Post-assessment reflection may help mitigate GenAI over-reliance, but this possibility requires empirical testing.
  • The findings extend feedback research by showing that students actively judge AI feedback within a relational context rather than receiving it passively.This selective trust is consistent with feedback literacy accounts emphasizing learners’ capacity to make judgments about feedback.
  • Disclosing ChatGPT’s role prompted students to examine fairness, objectivity, accuracy, authority, and the human instructor’s role rather than merely judging feedback content.Transparency did not produce disengagement; students combined agreement with feedback with critical evaluation of AI limitations and risks of over-reliance.
  • Post-assessment reflection shifted students from sentence-level improvement toward systemic questions about who should evaluate work and how tools, criteria, and human judgment interact.The inexpensive, technology-free activity surfaced reasoning that might otherwise remain invisible and could be embedded in courses using AI feedback.

6. Limitations and Future Work

The study’s conclusions are constrained by a small, homogeneous, potentially self-selected sample, possible response bias, and confounds in comparison with the earlier blinded study. Future work should test the proposed model through larger, more diverse, longitudinal, comparative, and independent research designs.

  • Sample limitations: The analytic sample included 13 reflective respondents from 19 enrolled students, creating a small and potentially self-selected dataset that limits generalizability.The authors note that the sample remains within ranges commonly considered adequate for reflexive thematic analysis of homogeneous, bounded samples.
  • Sample limitations: Transferability is restricted because all participants were male computing undergraduates from one course section, instructor, university, semester, and cultural context.The findings may not extend to women, non-computing students, other institutions, or different cultural contexts.
  • Researcher influence: Instructor-researcher involvement may have encouraged expected or favorable responses, so response bias cannot be entirely ruled out despite mitigation procedures and assurances about grades.The researcher served as both course instructor and principal investigator, creating a possible influence on students’ reflections.
  • Comparative limitations: Comparison with the earlier blinded study is confounded by different cohorts, academic years, and levels of cultural and institutional familiarity with GenAI.ChatGPT was relatively new in the earlier study, whereas GenAI tools had become normalized in higher education by the current study.
  • Future work: The Figure 6 model is a theoretical synthesis rather than a tested causal pathway, requiring longitudinal and comparative validation, especially of its instructor-authority feedback loop.Future research should use larger, more diverse samples; mixed-method or experimental comparison groups; multiple disciplines, institutions, and cultures; and researchers independent of instruction.

7. Conclusion

Within this exploratory study’s bounded sample and design, students critically engaged with disclosed AI assessment, valuing feedback for surface revision while reserving evaluative authority for human instructors. The conclusion proposes transparent, human-mediated, reflective GenAI integration while emphasizing tentative implications and the need for broader research.

  • Conclusion: Disclosure prompted critical engagement with evaluation, extending students’ attention from textual correction to authority, trust, fairness, and human teaching’s role.Participants also questioned the institutional purpose of higher education when assessment is delegated to AI.
  • Conclusion: Students found disclosed ChatGPT feedback clear, structured, and useful for surface-level revision but questioned AI’s appropriateness as the final evaluator.They particularly valued feedback on grammar, sentence structure, and organisation.
  • Conclusion: The proposed framework treats transparency, human mediation, and reflective practice as principles for using GenAI to augment rather than replace human judgment in assessment.Figure 7 presents this framework as derived from the exploratory study’s findings.
  • Conclusion: The implications are tentative because findings derive from a single course, semester, and thirteen participants, requiring study across other disciplines, institutions, and cultural settings.The conclusion presents open disclosure as a possible opportunity for reflection rather than as evidence of a universally effective approach.

Declaration Statements

The study distinguishes ChatGPT’s role as the research instrument being evaluated from Claude’s role as an analytic assistant, while making supporting materials available for reuse with identifying handwritten submissions withheld. Neither tool is credited with authorship, and the author retains responsibility for the work.

  • Data and materials: The reflective-task prompts, coding schedule, theme definitions, analysis log, and anonymised transcriptions are available in an Open Science Framework repository.Original handwritten submissions are withheld because handwriting may remain identifying after names are removed.
  • Use of AI: ChatGPT evaluated students’ scanned handwritten submissions, generating rubric-based feedback and numerical scores that constituted the study’s object of analysis.Students’ reflections on this AI-generated evaluation formed the dataset analyzed in the paper.
  • Use of AI: Claude Code conducted an independent confirmatory thematic-analysis pass on the 13 scanned handwritten reflections, supporting analytic refinement.The analysis followed Braun and Clarke’s (2006) inductive thematic-analysis approach and was documented in the methodology.
  • Authorship and responsibility: Neither AI tool is listed as an author, and the author retains responsibility for the work’s accuracy, originality, integrity, and conclusions.
Loading 2609.05346v1…