Source-linked AI summary

ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring

Shengjie Li, Vincent Ng

arXiv:2607.27671v1cs.CL

TL;DR

AES research has concentrated on ASAP, leaving uncertainty about generalization, fine-grained scoring, and valid cross-prompt evaluation. The paper introduces ICLE++, a corpus of persuasive essays with holistic and 10 trait-specific annotations, and reports that traits generally improve cross-prompt holistic scoring despite weaker trait-scoring results on ICLE++. The corpus is intended to support generalizability, multi-trait, and same-type cross-prompt AES evaluation, although its conclusions are limited to persuasive essays by university undergraduates who are non-native English speakers.

  • Problem

    ASAP-centered AES evaluation leaves generalizability and cross-prompt validity uncertain, while coarse or absent trait scores limit fine-grained scoring and feedback.

  • Method

    The paper introduces ICLE++, a corpus of persuasive essays from 10 prompts annotated with holistic and 10 trait-specific scores, and evaluates models with and without traits in within- and cross-prompt settings.

  • Results

    Traits generally improve cross-prompt holistic scoring, although trait scoring is poor on ICLE++ and trait inclusion has mixed effects in within-prompt scoring.

  • Takeaways & Limitations

    ICLE++ contributes an annotated corpus for testing AES generalizability and studying multi-trait and same-type cross-prompt scoring.

  • Takeaways & Limitations

    The findings are limited to persuasive essays written by university undergraduates who are non-native speakers of English and may not generalize to native-speaking high school students.

Abstract

from arXiv · show

The majority of the recently-developed models for automated essay scoring (AES) are evaluated solely on the ASAP corpus. However, ASAP is not without its limitations. For instance, it is not clear whether models trained on ASAP can generalize well when evaluated on other corpora. In light of these limitations, we introduce ICLE++, a corpus of persuasive student essays annotated with both holistic scores and trait-specific scores. Not only can ICLE++ be used to test the generalizability of AES models trained on ASAP, but it can also facilitate the evaluation of models developed for newer AES problems such as multi-trait scoring and cross-prompt scoring. We believe that ICLE++, which represents a culmination of our long-term effort in annotating the essays in the ICLE corpus, contributes to the set of much-needed annotated corpora for AES research.

1 Introduction

Recent AES research has relied heavily on ASAP, but its limited diversity and coarse trait structure leave generalizability, feedback, and cross-prompt evaluation unresolved. ICLE++ addresses these gaps with persuasive essays annotated for holistic and fine-grained trait scores.

  • Dataset limitations: Most recently developed AES models have been evaluated solely on ASAP, raising concerns about whether they generalize to other corpora.ASAP contains essays from U.S. students in grades 7–10, while other corpora may contain essays written by learners of English or different student populations.
  • Motivation: Holistic scores provide limited feedback because essay quality depends on multiple trait-specific dimensions, including organization and prompt adherence.Trait-specific scores could indicate which aspects of a low-scoring essay need improvement.
  • Fine-grained traits: Existing coarse-grained traits can restrict content-based feedback and prevent AES models from separately representing the traits considered by human scorers.This concern is especially relevant when content-based dimensions are grouped into a single trait.
  • Cross-prompt scoring: Within-prompt scoring may fail on new prompts without retraining, motivating cross-prompt scoring as a more challenging evaluation setting.Cross-prompt scoring trains on essays from some prompts and tests on essays written for unseen prompts.
  • Cross-prompt scoring: ASAP-based cross-prompt evaluation can mix persuasive, narrative, and source-dependent essays, making it unclear whether training and test essays share comparable scoring criteria.Different essay types may be judged using different rubrics, which complicates interpretation of cross-prompt results.
  • ICLE++ contribution: ICLE++ contains persuasive essays from 10 prompts annotated with holistic and trait-specific scores, complementing ASAP for generalizability and same-type cross-prompt evaluation.Its essays are written by university undergraduates from 16 countries, and most are between 500 and 600 words.

2 Corpus

ICLE++ is a persuasive-essay corpus with holistic and trait-specific annotations, designed to support analysis of AES generalizability, annotation quality, and trait-based scoring. Its comparisons with ASAP indicate that ICLE++ presents different feature relationships and new scoring challenges.

  • Corpus and annotation scheme: ICLE++ contains persuasive essays from university undergraduates in 16 countries, annotated for overall quality and 10 essay traits.The corpus was selected from the International Corpus of Learner English, and its overall-quality rubric uses scores from 1 to 4 in half-point increments.
  • Inter-annotator agreement: Inter-annotator agreement is substantial for OVERALL QUALITY and above 0.6 for every trait, but 31–40% of essays differ by exactly 0.5 points across traits.OVERALL QUALITY has α = 0.755; COHESION and ORGANIZATION are highest, while COHERENCE is lowest at 0.602. Larger disagreements are especially associated with PROMPT ADHERENCE and THESIS CLARITY.
  • Analysis of annotations: All traits correlate positively with OVERALL QUALITY, with ARGUMENT PERSUASIVENESS and DEVELOPMENT showing the largest feature weights and THESIS CLARITY the smallest.The reported correlations include DEVELOPMENT at 0.681, COHERENCE at 0.625, and THESIS CLARITY at 0.497; the feature-weight pattern follows these relationships.
  • Holistic and trait scoring: ICLE++ is more challenging than ASAP, and trait inclusion has mixed effects within prompts but generally improves cross-prompt holistic scoring.Gold traits produce much higher QWK scores, whereas predicted traits score poorly on ICLE++; nevertheless, traits slightly improve cross-prompt holistic scoring, possibly because PMAES is robust to noisy trait predictions.
  • Conclusion: Overall, the results suggest that ICLE++ presents new challenges for AES researchers.The corpus produces lower scoring results than ASAP and differs in feature relationships, annotation behavior, and the effects of trait information.

3 Conclusion

ICLE++ is a corpus of persuasive essays annotated with holistic and fine-grained trait-specific scores. The authors present it as a needed, publicly available resource for AES research.

  • ICLE++ contains persuasive essays annotated with both holistic scores and 10 fine-grained trait-specific scores.
  • The authors position ICLE++ as a valuable addition to the limited set of annotated corpora for automated essay scoring.
  • All ICLE++ annotations are made publicly available to AES researchers.

Limitations

The study’s findings are limited in scope because ICLE++ contains only persuasive essays written by university undergraduates who are non-native English speakers. Generalization to other essay types and student populations is therefore left unresolved by the authors.

  • The findings are limited to persuasive essays because the corpus focuses exclusively on that essay type.
  • ICLE++ consists of essays by university undergraduates who are non-native English speakers, so generalization to native-speaking high school students is unclear.

Ethics Statement

ICLE++ uses licensed source essays and releases annotations under research-oriented conditions. The dataset documentation also describes its prompts, score distributions, and trait rubrics.

  • Ethics Statement: Annotators were U.S.-based native English-speaking undergraduate students aged around 18–22 and were paid 10 US dollars per hour.
  • Ethics Statement: The source essays come from ICLE, so the dataset distributes annotations and identifiers rather than the essays themselves under non-profit research conditions.
  • Dataset statistics: ICLE++ contains 10 prompts, with documentation reporting essay length, language coverage, annotation counts, and average scores by prompt.
  • Scoring rubrics: Each essay trait is scored from 1 to 4 in half-point increments, where 4 denotes high quality and 1 denotes low quality.

C Analysis of the PMAES Features

The PMAES feature analysis compares feature correlations with Overall Quality across ICLE++ and ASAP. ICLE++ shows stronger readability correlations and more diverse prompt-level top features than ASAP.

  • Feature inventory: PMAES uses length-based, count-based, readability, essay-complexity, essay-variation, and other feature categories.
  • Correlation analysis: 0.24 versus 0.06: readability features average a stronger Pearson correlation with OVERALL QUALITY in ICLE++ than in ASAP.
  • Prompt-level analysis: ASAP’s prompt-level top features are usually length-related, whereas ICLE++ shows greater diversity among its strongest features.
  • Prompt-level analysis: The feature analysis provides suggestive evidence that constructing a high-performing AES system could be more challenging on ICLE++ than on ASAP.

D Additional Experimental Results

The paper reports additional experiments beyond QWK, including other evaluation metrics and per-prompt results.

  • D Additional Experimental Results: Additional experimental results are reported beyond the commonly used QWK metric.
  • D Additional Experimental Results: The additional results include metrics other than QWK.
  • D Additional Experimental Results: The additional results also include per-prompt analyses.

D.1 Results in terms of Other Metrics

Holistic scoring is evaluated on ASAP and ICLE++ using MAE, RMSE, and Pearson correlation. Agreement-based metrics generally indicate better AES performance on ASAP, while metric trends are not universally consistent.

  • Other metrics: Holistic scoring results are reported using mean absolute error, root mean squared error, and Pearson Correlation Coefficient.
  • Dataset comparison: AES systems generally perform better on ASAP according to agreement-based metrics such as QWK and Pearson Correlation Coefficient.
  • Metric trends: Within each dataset, higher QWK generally corresponds to higher Pearson correlation and lower MAE and RMSE.
  • Metric caveat: A cross-prompt comparison shows an exception in which Gold Traits has higher QWK but also higher MAE and RMSE than PMAES with traits.

D.2 Per-Prompt Results

This section reports per-prompt holistic-scoring results using QWK, MAE, RMSE, and Pearson correlation, while distinguishing prompt-level averaging from fold-level averaging. It also notes that gold-trait models perform best and introduces the models covered in the experiments.

  • D.2 Per-Prompt Results: Tables 30a and 30b report per-prompt holistic-scoring results on ASAP and ICLE++ using QWK.
  • D.2 Per-Prompt Results: The “Avg.” QWK values in these subtables differ from Table 9 because they macro-average per-prompt scores rather than cross-validation-fold scores.
  • D.2 Per-Prompt Results: Models trained on gold traits obtain the best results.
  • D.2 Per-Prompt Results: The highest average QWK model for each task-corpus combination does not always outperform every counterpart on every comparison.
  • D.2 Per-Prompt Results: Holistic-scoring results are also presented using MAE, RMSE, and Pearson Correlation Coefficient.
  • D.2 Per-Prompt Results: The section is part of the experiments’ per-prompt-results analysis, alongside the overview of models used and their implementation details.

E.1 Kumar et al.’s Model

This section presents Kumar et al.’s model and situates its evaluation within the reported ASAP and ICLE++ results. The model uses separate trait and holistic-score processing stacks with hierarchical essay representations.

  • E.1 Kumar et al.’s Model: Kumar et al.’s system is identified as a state-of-the-art model on the ASAP++ dataset.
  • E.1 Kumar et al.’s Model: The system uses a separate stack of layers for each trait score and the holistic score.
  • E.1 Kumar et al.’s Model: Each stack builds sentence representations with CNN-based processing and attention pooling, then forms document representations with an LSTM and attention pooling.

E.2 Uto et al.’s Model

This section describes Uto et al.’s BERT-based essay-scoring system and its extensions for trait prediction, alongside PMAES’s cross-prompt prompt-mapping approach. The models combine essay representations with handcrafted features for scoring.

  • E.2 Uto et al.’s Model: Uto et al.’s system concatenates a BERT essay embedding with handcrafted essay-level features before a linear scoring layer.
  • E.2 Uto et al.’s Model: Finetuning BERT with handcrafted features achieved state-of-the-art holistic essay-scoring performance on ASAP at the time.
  • E.2 Uto et al.’s Model: The Uto et al. simple variation predicts trait scores by expanding the final linear layer’s output neurons.
  • E.2 Uto et al.’s Model: The Uto et al. Kumar variation predicts traits first and then uses trait scores with essay embeddings for holistic-score prediction.
  • E.2 Uto et al.’s Model: PMAES addresses cross-prompt scoring by mapping source and target prompts and aligning their essay representations with contrastive learning.
  • E.2 Uto et al.’s Model: PMAES predicts holistic scores with linear layers over essay representations concatenated with handcrafted features.

E.4 Implementation Details

The implementation varies training schedules, learning rates, optimizers, and hardware across Kumar, Uto, PMAES, and the gold-trait linear regressor. Hyperparameter tuning uses development data with fixed random seed 11.

  • E.4 Implementation Details: All models tune learning rate and dropout rate on development data across the listed search values and use random seed 11.
  • E.4 Implementation Details: Kumar’s model trains for 150 epochs on ICLE++ and 100 on ASAP using AdamW, batch size 64, and a single RTX 3090.
  • E.4 Implementation Details: Kumar, Uto et al., and PMAES require approximately 4, 4, and 5 hours of training, respectively.
  • E.4 Implementation Details: Uto et al.’s models train for 50 epochs on ICLE++ and 20 on ASAP using a 6 × 10^-4 learning rate and a single RTX A6000.
  • E.4 Implementation Details: PMAES trains for 50 epochs on ICLE++ and 20 on ASAP with a 3 × 10^-4 learning rate and Adam.
  • E.4 Implementation Details: The gold-trait linear regressor uses scikit-learn with default hyperparameter values.
Loading 2607.27671v1…