Source-linked AI summary
A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment
María Eugenia Curi, Germán Capdehourat, Isabel Amigo, Magdalena Romano, Rosana Serra, Adrián Silveira, Andrés Peri
TL;DR
Large-scale evidence for reliable LLM-assisted scoring of high-stakes Spanish writing remains limited, while manual writing assessment is resource-intensive. This paper validates an AI-assisted framework using operational national-assessment data and a human-in-the-loop decision flow, finding stable alignment with human evaluation and at least 50% potential reduction in full human scoring while preserving decision quality.
Problem
Evidence remains limited for reliable, valid, and appropriately supervised LLM scoring in real large-scale operational contexts, especially high-stakes Spanish writing assessment.
Method
The study evaluates prompt-based LLM scoring across two national-test editions and integrates it into a human-in-the-loop workflow that examines rubric agreement, decisions, and workload.
Results
At least 50% of written responses could avoid full human scoring while preserving decision quality, and cross-year validation suggests stable generalization with minimal prompting adjustments.
Takeaways & Limitations
Carefully designed human oversight can support scalable, timely, and consistent large-scale writing assessment while retaining expert review for critical cases.
Takeaways & Limitations
The analysis uses human-rater agreement as ground truth despite non-negligible rater variability, and operational deployment may introduce monitoring, drift, and edge-case challenges.
Abstract
from arXiv · showhide
The integration of artificial intelligence (AI), particularly large language models (LLMs), into educational assessment has opened new opportunities to enhance the efficiency and scalability of grading processes. This study presents the design and validation of an AI-assisted scoring framework for written responses in a large-scale national assessment. The proposed approach focuses on short written texts of approximately 150-200 words and incorporates a human-in-the-loop strategy to preserve assessment quality while reducing manual workload. The study is grounded in a real operational context, using data from two recent editions of a nationwide test, each comprising approximately 5,000 student responses. We analyze the alignment between AI-generated scores and human raters across multiple rubric dimensions, as well as the impact of the proposed decision flow on pass/fail outcomes. Results show moderate to high agreement between the model and human evaluations in most dimensions, supporting the feasibility of AI assistance in this setting. Moreover, the proposed correction workflow identifies cases where human review is most valuable, enabling a more efficient allocation of expert effort. The findings suggest that AI-assisted scoring can be safely integrated into large-scale assessment processes only when combined with carefully designed human oversight. The paper concludes by discussing practical implications for deployment in national assessment systems and outlining future research directions, including longitudinal monitoring of model-human alignment and the analysis of potential cognitive bias introduced by AI-supported review workflows.
1 Introduction
This study evaluates LLM-based scoring for Spanish argumentative writing in a national high-stakes assessment and designs a human-in-the-loop workflow to support reliable, scalable decisions. It examines agreement with human raters, cross-edition behavior, pass/fail implications, and potential workload reduction.
- Motivation: Evidence from real large-scale operational contexts remains limited, particularly regarding reliability, validity, and the appropriate role of human oversight.
- Assessment context: The study analyzes Spanish argumentative texts of approximately 150-200 words scored with a detailed analytic rubric by trained human raters.
- Assessment context: Operational data from the 2024 and 2025 national tests comprise approximately 5,000–6,000 participants per edition, enabling realistic cross-year analysis.
- Study objectives: The evaluation covers AI-human agreement, proficiency classification, pass–fail outcomes, cross-edition generalization, consistency, and potential workload reduction.
- Framework: The proposed HITL framework combines automated scoring with expert review to accelerate reporting, optimize expertise, and maintain high-stakes assessment quality.
- Contributions: The paper contributes empirical evidence on operational LLM integration and outlines a feasible pathway toward responsible, scalable AI-assisted evaluation.
2 Literature Review and Research Positioning
Prior work establishes a broad metric and application landscape for automated grading, but evidence for large-scale Spanish argumentative-writing assessment remains sparse. This study positions its contribution around operational integration, human oversight, and decision-oriented validation.
- Existing literature: Automated grading research spans short answers, open-ended questions, essays, computer science, science, chemistry, and mathematics, but most text-based studies use English data.
- Spanish-writing gap: Research on AI-assisted Spanish writing grading remains limited, with few studies addressing rubric-based argumentative writing in large-scale examinations.
- Research positioning: Because findings across use cases are difficult to generalize, prior results provide limited direct evidence for high-stakes Spanish-language assessment.
- Evaluation metrics: Validation metrics include agreement measures, classification metrics, uncertainty indices, robustness, explainability, and time saved by automation.
- Evaluation metrics: The literature emphasizes that uncertainty metrics are setting-dependent and that human oversight is important for transparency and trustworthy evaluation.
- Evaluation metrics: AUC evaluates whether incorrect responses receive higher uncertainty than correct responses, while C-Index evaluates whether larger true errors receive higher uncertainty scores.
3 Research Gap and Contribution
The paper addresses a gap in safely integrating LLM scoring into operational, high-stakes assessment workflows. Its contribution combines decision-oriented evaluation with real Spanish data and explicit human supervision.
- Research gap: Operational assessment research has devoted comparatively little attention to safely integrating grading models into workflows where outcomes determine certification decisions.
- Contribution: The study evaluates AI not only through agreement metrics but also by examining how predictions affect final decisions and which cases require human review.
- Contribution: Real-world Spanish data and expert validation address the limited transferability of findings from English-based assessment studies.
- Contribution: The proposed prompt-engineering system integrates Spanish argumentative-writing scoring into a high-stakes national examination and a human-in-the-loop workflow.
- Contribution: The framework is designed so that final decisions affecting test outcomes receive appropriate human oversight while reducing expert grading work.
4 Context and Case Study
The case study examines AI-assisted scoring within Uruguay’s Acredita EB national certification test, whose manually scored Writing section is the main source of evaluation workload. The assessment uses a multi-stage human scoring process and a detailed rubric spanning discursive, textual, and orthographic domains.
- 4 Context and Case Study: Acredita EB is Uruguay’s annual national accreditation test for adults seeking lower-secondary certification.The test is organized by ANEP and has accumulated substantial human-scored data since 2020.
- 4.1 Structure of the Exam: The computer-based assessment comprises Reading Comprehension, Problem Solving, and Writing sections completed in one 2-hour-50-minute session.The sections are graded independently by separate technical teams and examiners.
- 4.1 Structure of the Exam: Writing requires a short argumentative text demonstrating a clear position, supporting arguments, organization, cohesion, and language control.The annual topic varies across environmental, ethical-social, and related issues.
- 4.1 Structure of the Exam: Manual Writing evaluation is the most time- and resource-intensive scoring component, delaying final results to no sooner than three months after testing.The other two sections are primarily multiple-choice and scored automatically.
- 4.2 Evaluation Procedure: The Writing workflow progresses from calibration, to double-scoring for agreement checks, to single-rater scoring with sampled expert review.After section scoring, IRT and bookmark methods determine proficiency levels, which feed the final pass/fail decision.
- 4.3 Writing Rubric: The Writing rubric organizes 15 scoring items across discursive, textual, and orthographic domains.The discursive domain covers communicative purpose, organization, register, and vocabulary; the textual domain covers grammatical mechanisms supporting textual unity.
5 Methodology
The methodology evaluates prompt-based LLM scoring against operational human scores using data from the 2024 and 2025 test editions. It combines item-level agreement analysis, cross-year validation, and workflow-oriented assessment of AI assistance.
- 5 Methodology: The dataset contains participant proficiency levels, Writing responses, and item-level human scores from the 2024 and 2025 test editions.Calibration texts additionally include independent scores from language experts and consensus scores.
- 5.1 Dataset Description: Ten evaluators independently scored 50 calibration texts in 2024, enabling an empirical upper bound for expected AI performance.The calibration analysis reports minimum, average, and maximum agreement for each rubric item.
- 5.1 Dataset Description: No rubric item achieved complete evaluator agreement; minimum agreement was generally below 80%, while maximum agreement exceeded 90% for several items.Other items had maximum agreement closer to 80%.
- 5.1 Dataset Description: Cohen’s Kappa indicated substantial or moderate agreement for more than half of the rubric items, with lower values partly explained by skewed score distributions.Vocabulary and Thematic progression showed slight agreement, while Opinion, Argumentation, and Connectors showed fair agreement.
- 5.2 AI Evaluation: AI performance was evaluated against operational human scores using random samples of 1,000 responses from each test edition that were excluded from prompt refinement.These single-rater operational scores serve as the complete-dataset reference labels despite possible inconsistencies.
- 5.2 AI Evaluation: The model uses dedicated Spanish prompts for each rubric item to assign scores, identify supporting excerpts, and justify decisions.Prompts were iteratively tested on 10 texts, then applied to the evaluation samples; outputs were compared using exact-match accuracy at item and aggregate levels.
- 5.3 Cross-Year Validation: Cross-year validation tests 2025 responses with prompts and contextual patterns derived from 2024 without retraining.This procedure assesses robustness and prompt generalization across exam editions.
6 Results
The AI model was evaluated for accuracy, consistency, cross-year generalization, rubric-level agreement, and effects on proficiency and pass–fail decisions. Results indicate stable performance, systematic conservatism, and a need to review AI-predicted failures.
- Calibration analysis: The 2024 calibration analysis evaluated AI accuracy and output consistency against human scoring across rubric items.Each calibration text was scored repeatedly to assess stability during model development.
- Calibration analysis: For most rubric items, consistency approached or exceeded 90%, meaning the model repeated the same score in roughly 9 of 10 evaluations.Spelling was the main exception and was assigned to a deterministic LanguageTool-based procedure.
- 2024 large-scale evaluation: 60%–80% accuracy was observed for most rubric items in the unseen 2024 sample, with vocabulary, syntax, and spelling among the lowest-performing items.The lower performance was linked to rubric distinctions that are difficult for token-level evaluation to capture.
- Cross-year generalization: Cross-year performance variations were very small, indicating that prompts developed for 2024 generalized to the 2025 assessment without major adjustments.Only minor prompt changes were made for rubric items tied explicitly to the annual writing-task instructions.
- Cross-year generalization: Most 2025 rubric items showed moderate or fair Cohen’s Kappa agreement, below the human calibration agreement.Moderate agreement was reported for Introduction, Conclusion, Register, Nominal agreement, and Subject–verb agreement; fair agreement was reported for several other items.
- Proficiency and pass–fail decisions: The AI model applied identical cut-score criteria to AI and human scores, but its proficiency classifications were systematically stricter.The confusion-matrix analysis indicates predictable under-grading rather than random error, with higher agreement for Proficient than Insufficient responses.
- Proficiency and pass–fail decisions: 15.3% in 2024 and 16.5% in 2025 of human-passing tests were classified as failing under AI-based scoring, supporting human review of AI-predicted failures.The overall pass–fail analysis combined Writing results with Reading Comprehension and Problem Solving outcomes.
7 Human-in-the-Loop AI-Assisted Evaluation
The proposed framework combines automated Writing scoring with targeted expert review in a high-stakes assessment workflow. It is intended to preserve decision quality while reducing the amount of full human scoring required.
- Framework rationale: The study proposes a Human-in-the-Loop framework that combines AI-assisted Writing scoring with expert human review.The framework is designed to retain quality and consistency with human evaluation while supporting more efficient result processing.
- Operational workflow: The workflow begins with calibration to validate the model, adapt prompts to the current topic, and identify issues before automated scoring.The proposed calibration stage could include 50 to 100 written productions.
- Operational workflow: Automated scoring assigns rubric-level scores and proficiency levels, after which human raters review candidates whose pass–fail outcome depends on Writing.IRT and the Bookmark method are applied to establish cut scores before targeted review.
- Decision safeguards: Candidates with Proficient results across all three sections pass, while human review can address cases in which the AI Writing score may alter the final outcome.The workflow includes recalibration after review to account for possible shifts in proficiency levels and decisions.
- Expected workload reduction: At least 50% fewer written productions would require full human scoring under the proposed process.The exact reduction depends partly on candidate performance in the other test sections and may vary across assessment editions.
8 Conclusion and Future Directions
The study finds that prompt-based AI scoring can align acceptably with human evaluation and reduce manual scoring through a Human-in-the-Loop workflow, while requiring further operational validation and monitoring.
- Prompt-based AI scoring aligned with an existing analytic rubric and human evaluation workflow using operational data from the 2024 and 2025 test editions.
- The model showed variable item-level agreement across rubric criteria but stable, systematically conservative behavior and acceptable alignment in proficiency-level and pass–fail decisions.
- The proposed HITL framework could reduce by at least 50% the responses requiring full human scoring while preserving decision quality and expert oversight in critical cases.
- The study’s offline experimental conditions leave deployment challenges involving monitoring, model drift, and edge cases unresolved.
- A controlled pilot in future test editions and continued longitudinal monitoring are proposed to assess operational impact and detect performance drift.
- Future research should examine whether exposure to AI-generated scores introduces anchoring or automation bias in human raters’ judgments.
- Further work should refine rubric-sensitive prompting for vocabulary, syntax, and punctuation and explore hybrid LLM–NLP approaches.
- The findings provide empirical evidence that carefully designed HITL systems can support large-scale writing assessment processes.
Declarations
The study used anonymized secondary assessment data supplied by Uruguay’s National Public Education Administration and reports no specific external grant funding.
- The research received no specific grant from public, commercial, or not-for-profit funding agencies.
- The dataset consisted of fully anonymized written responses and corresponding human-assigned scores, with no personally identifiable information.
- The assessment data were provided by Uruguay’s National Public Education Administration, ANEP.
- LLM evaluations used the OpenAI Enterprise API under terms stating that submitted data were not retained for model training or future model improvement.