Source-linked AI summary

Language Proficiency Assessment from Eye Movements in Naturalistic Passage Reading

Shachar Frenkel, Ido Falah, Omer Shubi, Yevgeni Berzak

arXiv:2608.30583v1cs.CL

TL;DR

Standard proficiency tests are costly and rely on limited offline responses, while eye movement assessment could capture proficiency during ordinary reading. This paper validates and extends that approach across passages, measures, models, and reading regimes, finding effectiveness alongside L1 bias and greater reliability than standard scores.

  • Problem

    Evidence for eye movement based proficiency testing remained limited beyond single-sentence data, with unresolved questions about L1 bias and score reliability.

  • Method

    The study evaluates eye movement based proficiency assessment across contextualized passages, proficiency measures, prediction models, and comprehension and information-seeking reading regimes.

  • Results

    The approach is effective across the evaluated settings; EyeScore favors L1s closer to English, while debiasing alleviates this issue and reliability exceeds standard test scores.

  • Takeaways & Limitations

    The findings provide substantial support for the viability of eye movement based proficiency testing as a standalone technology or an add-on to traditional tests.

  • Takeaways & Limitations

    Eye tracking based assessment currently covers comprehension-related linguistic abilities, leaving spoken and written production abilities out of scope.

Abstract

from arXiv · show

Standard language proficiency tests rely on linguistic tasks such as vocabulary, grammar and reading comprehension quizzes. An alternative, cognitively motivated approach, introduced in Berzak et al. (2018), proposed instead to predict language proficiency from behavioral traces of eye movements in reading. In this work, we validate and extend this approach from single sentences to more naturalistic reading of contextualized passages in English as a second language, new proficiency measures, prediction models, and reading in an information seeking regime. We find that the approach is effective in all these evaluations. We further address two key open questions on eye movement based proficiency testing: (1) potential scoring biases that reflect the proximity of the reader's native language to English, which may undermine validity, and (2) its reliability. We find that eye movement based proficiency scores are indeed biased towards L1s that are linguistically closer to English. We propose a score debiasing method which effectively remedies this issue. The reliability analyses suggest that eye movement proficiency scores are more reliable than standard language proficiency scores. Overall, our results strengthen and broaden the empirical foundations for future eye movement based language assessment technologies.

1 Introduction

This work extends eye movement based English proficiency assessment to naturalistic passage reading and examines its validity, bias, and reliability. It finds effective assessment across broader settings, while identifying L1 bias and proposing debiasing.

  • Standard assessments are costly and labor-intensive, rely on small sets of offline responses, and provide limited access to online language-processing signals.These constraints affect test development, deployment, and feedback quality.
  • Eye movement based assessment derives L2 proficiency from behavioral traces during ordinary reading rather than dedicated test items.The approach can use eye movements as online cognitive signals and may avoid several practical drawbacks of traditional testing.
  • The study validates the approach on 1,265 L2 participants across two passage-reading corpora, multiple proficiency measures, prediction models, and comprehension and information-seeking regimes.
  • EyeScore favors L2 speakers whose native languages are more similar to English, and the authors propose a debiasing method that alleviates this issue.
  • EyeScore is analyzed for internal consistency and split-half reliability, with results indicating greater reliability than standard language proficiency test scores.

2 Background and Related Work

The paper grounds eye movement proficiency prediction in psycholinguistic accounts of online reading and L1 effects on L2 processing. It extends earlier single-sentence work within a broader research context on reader characteristics and reading regimes.

  • Eye movements reflect real-time language comprehension processes, while proficiency differences affect how readers process linguistic input.
  • The proposed causal chain links proficiency differences to online comprehension processes and then to behavioral traces in reading eye movements.
  • The study builds on Berzak et al. (2018), extending feasibility evidence obtained from single-sentence reading with external proficiency references.
  • Its L1-validity analysis is motivated by evidence that native language influences L2 learning, production, comprehension, and reading.
  • The work belongs to an emerging line of research predicting reader characteristics and reader-text interactions from eye movements, including dyslexia and reading regimes.

3 Eye Movement Based Estimation of Language Proficiency

The paper evaluates two eye movement approaches to L2 proficiency: standalone EyeScore and prediction of external proficiency-test outcomes. EyeScore compares normalized L2 eye movement patterns with an L1 prototype, while prediction models include regression and machine-learning baselines.

  • EyeScore: EyeScore estimates L2 proficiency from the cosine similarity between a reader’s normalized eye movement vector and an averaged normalized L1 prototype.Participant eye movement features are z-scored using an L2-fitted scaler before comparison.
  • External-test prediction: The external-test approach uses eye movement feature vectors to predict standardized proficiency scores without administering the target test.The evaluated models include Ridge regression, LightGBM, and TabPFN-3.
  • Baselines: Reading speed, measured in words per minute, is the primary baseline for estimating the added value of eye tracking.WPM can be obtained without eye tracking equipment.
  • Baselines: Average Train Score assigns each test participant the mean external-test score of the training set.

4 Eye Movement Representations

Eye movement representations combine fixation, saccade, and word-property information. Some feature sets require identical texts across participants, whereas others support participants reading different texts.

  • Eye movement trajectories are represented through fixations, stable gaze periods, and saccades, rapid transitions between fixations.
  • Average Fixation Metrics aggregate fixation and saccade measures across words, including durations, skips, and regressions.
  • Syntactic Clusters average fixation measures by part-of-speech categories using Universal and Penn Treebank tags.
  • Word Property Coefficients encode how word length, frequency, and surprisal influence each reader’s fixation measures through linear-model coefficients.Surprisal is obtained from GPT-2 and frequency from Wordfreq.
  • Transitions record saccade counts between word pairs by direction, while Word Fixation Metrics retain gaze and total fixation durations for each word token.
  • Transitions and Word Fixation Metrics require fixed texts across participants; the other three feature sets also work in the Any Text regime.

5 Experimental Setup

The study evaluates eye movement based English proficiency assessment on large passage-reading datasets, using multiple proficiency measures, prediction models, and reading regimes. It compares Fixed Text and Any Text settings and evaluates standalone EyeScore alongside prediction of external test scores.

  • Datasets: MECO L2 contains 12 informational passages, whereas OneStopL2 contains 30 newswire articles presented in original or simplified form.MECO L2 participants read the same texts; OneStopL2 participants read one of six batch-by-difficulty subgroups, with comprehension or information-seeking instructions.
  • Proficiency measures: The evaluation uses LexTALE, Michigan Placement Test, and a MECO L2 Composite Proficiency score derived from four tests with PCA.LexTALE measures vocabulary, Michigan includes listening, grammar, vocabulary, and reading comprehension, and the Composite combines spelling, vocabulary, and two TOWRE variants.
  • Evaluation regimes: Fixed Text uses eye movement data from the same texts as the test participant, while Any Text uses data from different texts and supports arbitrary test materials.The study assumes no prior eye movement data for the test participant; Fixed Text is analogous to specific test forms, whereas Any Text is more general.
  • Evaluation measures: EyeScore is evaluated by Pearson correlation with external proficiency tests, while direct prediction uses Pearson r and Mean Absolute Error between predicted and true scores.Both approaches use cross-validation, with L2 participants in the test set; external-score prediction includes L1 participants in training.
  • Evaluation procedure: The cross-validation procedures use leave-one-participant-out or fold-based splits, with Ridge regularization tuned on training data across 10 log-spaced values from 10^-3 to 10^4.MECO L2 and OneStopL2 use different participant and passage splits for Fixed Text and Any Text evaluations.

6 Replication Results

Eye movement based proficiency assessment generalizes from single sentences to naturalistic passage reading across datasets, reference tests, models, and reading regimes. EyeScore is robust across textual-input conditions, while direct prediction generally performs better but depends on training-data size.

  • EyeScore: EyeScore correlations are consistent across two datasets and three external tests, with mostly similar or higher correlations than Berzak et al. (2018).Among feature sets applicable to both regimes, WP-Coefs tends to perform best; Word Fixation Metrics is strongest in Fixed Text.
  • EyeScore: Comparable EyeScore results in Fixed Text and Any Text indicate robustness to the specifics of the textual input.Similar results also appear in the OneStopL2 information-seeking regime, supporting robustness across reading manner.
  • EyeScore: Eye movement feature sets systematically outperform the WPM reading-speed baseline, despite WPM performing much more strongly here than in Berzak et al. (2018), where r ≤0.28.The authors hypothesize that strategic rereading was more feasible in the earlier sentence-reading task than in passage reading.
  • External-score prediction: Direct prediction of external proficiency scores generally yields higher Pearson r and larger improvements over WPM than EyeScore.The primary exception is OneStopL2 Fixed Text, where Michigan and LexTALE prediction is low, which the authors attribute to small training data.
  • External-score prediction: External-score prediction largely generalizes to information seeking, while LightGBM matches Ridge on MECO L2 but tends to perform worse on OneStopL2.TabPFN-3 largely outperforms Ridge on MECO L2 and performs similarly on OneStopL2.
  • Benchmarks: Pearson r values of 0.55 between LexTALE and Composite in MECO L2 and 0.61 between LexTALE and Michigan in OneStopL2 provide dataset-specific benchmarks for EyeScore and prediction.Correlations among widely used standardized English proficiency tests are typically reported in the range 0.61–0.85.

7 Measuring and Mitigating L1 Bias in EyeScore

EyeScore is biased toward L2 readers whose native languages are linguistically closer to English, but a regression-based correction substantially mitigates this bias across evaluation settings.

  • Measuring L1 Bias: EyeScore residuals are regressed on L1–English distance after controlling for an external proficiency test to detect proximity bias.The control test may be LexTALE, Composite, or Michigan; a significantly negative distance coefficient indicates bias toward L1s closer to English.
  • Measuring L1 Bias: −0.52 βdist, r = −0.18, p < 10^-7: participants with L1s less similar to English scored lower than their LexTALE reference predicted.The negative relationship also appears across feature sets and several dataset–test combinations, although some coefficients are nonsignificant.
  • Debiasing EyeScore: EyeScoredb subtracts the L1-dependent bias predicted from the participant’s L1 distance to English.The correction uses a model fitted relative to an external proficiency test and can be applied with Seen L1 or held-out New L1 splits.
  • Debiasing EyeScore: In MECO L2, debiased distance coefficients were significantly smaller than the original EyeScore coefficients in all evaluated cases.The strongest effectiveness appears for Seen L1 splits, while New L1 splits provide a stricter generalization test.
  • Debiasing EyeScore: Debiasing nearly eliminated LexTALE-related bias and also removed much of the bias when evaluated against Composite and Michigan scores.Similar mitigation patterns held for Fixed Text and information-seeking reading, supporting cross-setting robustness.
  • Debiasing EyeScore: Average correlation changes after debiasing were small, while the average EyeScore–EyeScoredb correlation was 0.99.Across Any Text evaluations, mean Δrdb ranged from −0.016 to +0.006 depending on dataset and reference test.

8 EyeScore Reliability

EyeScore reliability is assessed through internal consistency and split-half reliability, producing very high values that generally exceed those reported for standard English proficiency tests.

  • Measures: Reliability is evaluated using Cronbach’s α for internal consistency and split-half Pearson r across passage partitions.Split-half estimates average Pearson correlations over 20 random passage splits.
  • Measures: Cronbach’s α requires multiple item-level scores, so EyeScore is aggregated per passage for the internal-consistency analysis.In MECO L2, the analysis uses 144 participants with eye-movement data for all 12 passages.
  • Results: EyeScore averages Cronbach’s α = 0.98 and split-half Pearson r = 0.93 in the Any Text regime.The results are described as extremely high in nearly all cases.
  • Results: Similar reliability results are obtained for Fixed Text and information-seeking reading.These analyses extend the Any Text findings across alternative reading regimes.
  • Results: EyeScore reliability tends to exceed standardized-test reliability: Michigan has α = 0.91 and split-half r = 0.92 in OneStopL2.Reported comparison values include α = 0.91 for IELTS reading and listening sections and overall TOEFL-iBT reliability of 0.90.

9 Conclusion

The study replicates and broadens eye-movement-based language proficiency testing from single sentences to passage reading across datasets, proficiency references, and reading regimes. It also finds support for the method’s validity and reliability while identifying L1 bias as a correctable issue.

  • Conclusion: This is the first replication study of deriving language proficiency from eye movements during reading instead of standard test items.The study evaluates the paradigm across passage-reading datasets, proficiency references, and reading regimes.
  • Conclusion: The methodology is effective across datasets, standard proficiency references, and reading regimes, while its validity and reliability are also examined.The conclusion includes investigation of L1 bias alongside reliability analysis.
  • Conclusion: The experiments provide substantial support for the viability of eye-movement-based proficiency testing.The authors envision standalone deployment or use as an add-on to traditional proficiency tests.

10 Limitations

The study’s limitations concern the scope of eye-tracking assessment, assumptions in L1-bias correction, dataset coverage, laboratory conditions, and untested behaviors and reliability across sessions.

  • Eye-tracking assessment covers comprehension-related linguistic abilities but not spoken or written production abilities.
  • L1 debiasing depends on language-distance estimates, an external reference test, unbiased external tests, and participants having a single L1.If external tests are L1-biased, the procedure underestimates the bias to remove.
  • The datasets are restricted to English, two textual domains, specific language backgrounds, and a limited age range.Broader evaluation would require additional languages, domains, populations, and native-language backgrounds.
  • The datasets were collected with state-of-the-art eye tracking equipment in laboratory conditions, leaving performance with lower-quality equipment and outside laboratories to future work.
  • Further experimentation is needed to characterize cheating through strategies such as skimming and test reliability across different sessions.

11 Ethical Considerations

The supplied passages describe dataset governance and the eye-tracking measures and feature sets used in the study, including anonymization, consent, and institutional oversight.

  • The three eye-tracking datasets were collected under institutional IRB protocols, with written participant consent and anonymized data.
  • The datasets were collected for studying the relation between eye movements and linguistic proficiency, while anonymization precludes reader-identification uses.
  • The feature inventory includes first fixation, gaze, total fixation, go-past, skip-rate, and regression-probability measures.The first four are averaged across words, while skip rate and regression probability are computed over presented textual materials.
  • S-Clusters contains 258 part-of-speech-conditioned features, while WP-Coefs contains 24 participant- and measure-specific linear-model coefficients.
  • Transition features count saccades between word pairs, and Word Fixation Metrics contain thousands of word-level features subject to a 2,000-word TabPFN-3 limit.
  • Compared with Berzak et al. (2018), the study uses raw rather than speed-normalized measures and adds go-past, skip, regression, and regression-prediction features.

B.1 Eye Tracking Datasets

The study uses three English eye-tracking datasets and evaluates proficiency scoring, prediction, L1-bias correction, and reliability across reading regimes and models.

  • The study uses MECO L2 v2.1, OneStop v1.0.3, and OneStopL2 v0.1 eye-tracking datasets.
  • OneStopL2’s first reading portion contains 1,600,086 word tokens and 2,677,486 fixations, while OneStop contains 2,097,963 tokens and 2,027,969 fixations.
  • EyeScore is evaluated against standard proficiency tests, while external-test prediction uses Ridge regression, LightGBM, and pretrained TabPFN-3 models.
  • OneStopL2 includes reading-for-comprehension and information-seeking regimes, with 20 L2 and 20 L1 training participants in Fixed Text and 81 L2 and 81 L1 in Any Text.
  • EyeScoredb adjusts EyeScore using a cross-validated c_resp,Test model under Seen L1 and stricter New L1 splits.
  • Reliability measures include stratified Cronbach’s α, split-half Pearson r, and standard error of measurement; reported overall α values are 0.91 and 0.95 in two regimes.
Loading 2608.30583v1…