Source-linked AI summary
Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis
Alona Strugatski, Licol Zeinfeld, Giora Alexandron
TL;DR
Human-targeted assessments are widely used to infer LLM abilities, but such inferences require comparable latent structures. Using EFA, factor congruence, and resampling across chemistry and quantitative reasoning, the study finds systematic human–LLM structural differences that raise validity concerns.
Problem
Using human-designed assessments to infer LLM abilities assumes that human-validated links between performance and constructs also hold for LLMs.
Method
The study compares two human assessment datasets with responses from six multimodal LLMs using EFA, factor congruence, and repeated resampling.
Results
Across both instruments, LLM-human latent-structure similarity was lower than human-human similarity, with systematic factor-structure differences.
Takeaways & Limitations
The analyzed assessments may measure different constructs for humans and LLMs, raising validity concerns about using them for AI-capability claims.
Takeaways & Limitations
External validity is constrained by the small number of instruments and specific GenAI tools, while LLM response patterns may vary with prompting and model architecture.
Abstract
from arXiv · showhide
The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skills and competencies in humans, such as standardized exams, and to use performance on these instruments as evidence for generalizable claims about LLMs' underlying abilities on the same skills the assessments are intended to measure in humans. However, from a validity perspective, such inferences require that the relationship between observed performance and underlying constructs established for humans also holds for LLMs. In particular, a necessary condition for transferring score interpretations is similarity in the latent structure of responses to the assessment. In this study, we examine whether this condition holds in two educational contexts: high-school chemistry and a quantitative reasoning section of a university entrance exam. Using a case study design, we compare human response data with responses generated by six multimodal LLMs. Our analytical approach combines exploratory factor analysis, factor congruence, and resampling to assess latent structure similarity across human learners and LLMs. Across both instruments, we find systematic differences between human and LLM factor structures, showing evidence that the analyzed assessments may not measure the same constructs for humans and LLMs. These findings call into question the validity of evaluation practices that use educational assessments to make claims about AI capabilities.
1 Introduction
The study asks whether assessments designed for humans support the same construct-level interpretations for LLMs. It examines latent-structure similarity across chemistry and quantitative-reasoning assessments using EFA, factor congruence, and resampling.
- Human-validated links between test performance and underlying abilities may not extend unchanged to LLMs.
- Similar latent structures are a necessary condition for transferring the same score interpretation across human and LLM populations.
- The research question concerns whether science-education and mathematical-reasoning assessments exhibit similar latent structures for humans and LLMs.
- The case study compares a high-school chemistry test and university quantitative-reasoning exam responses from humans and six multimodal LLMs.
- Separate human and LLM EFAs were followed by factor congruence, optimal matching, and repeated resampling to quantify structural similarity.
- Across two case studies, established assessments may capture substantially different constructs in humans and LLMs.
2 Related Work
Related work identifies growing use of human-targeted assessments for LLM evaluation alongside concerns about validity, response patterns, reasoning brittleness, and benchmark transparency.
- Human-targeted assessments are increasingly reused to evaluate LLM knowledge and reasoning across medicine, mathematics, and science.
- High scores on some clinical benchmarks may fail to translate to clinical decisions requiring the same knowledge.
- Benchmarking concerns include data contamination, brittle reasoning under minor perturbations, aggregate metrics, and limited instance-level auditability.
- LLM responses often differ from human examinees and may lack consistent psychometrically plausible profiles.
- STEM examples report limitations on diagram-based engineering questions and under-specified physics problems.
3 Methodology
The methodology compares human and pooled LLM latent structures across two multimodal assessments using separate EFAs, factor-retention criteria, factor matching, and repeated resampling.
- The pipeline preprocesses responses, estimates retained factors, fits separate human and LLM EFAs, and compares structures quantitatively.
- The instruments comprise a 22-item chemistry test with 931 students and a 20-item quantitative-reasoning exam with over 4,800 examinees.
- Non-public human data were chosen to preserve authentic learner effort, data quality, and limited prior exposure to the instruments.
- Six multimodal LLMs produced 120 pooled responses per instrument through web interfaces using new temporary chat sessions.
- Binary responses were analyzed with tetrachoric correlations, while factor retention used the Kaiser criterion and parallel analysis.
- Factor-loading matrices were compared with cosine-based congruence, Hungarian matching, mean absolute congruence, and 100 resampled iterations.
4 Results
Human and LLM factor structures differed across both assessments, despite some agreement in retained factor counts. Resampled human-human similarity consistently exceeded LLM-human similarity.
- Factor retention: Parallel analysis retained five factors for Chemistry humans versus most often four for LLMs, and 7–8 versus five for Quantitative Reasoning.
- EFA structures: Under Kaiser retention, Chemistry had four factors and Quantitative Reasoning five for both groups, but item-loading structures still differed.
- EFA structures: In Quantitative Reasoning, one LLM factor loaded above 0.5 on five items that humans distributed across three factors.
- Congruence: Human-human congruence distributions were shifted higher than LLM-human distributions, indicating greater structural similarity among independently resampled humans.
- Congruence: The human-human baseline was variable, so the key result was the consistent gap separating human-human from LLM-human matching.
- Figure 1: Figure 1 compares blue human and red LLM EFA structures across Chemistry and Quantitative Reasoning under Kaiser-retained solutions.
5 Discussion
Across both assessment instruments, qualitative and quantitative analyses found considerable differences between human and LLM latent factor structures. These differences create a validity-of-inference concern for interpreting benchmark scores as evidence of the same underlying abilities.
- LLM–human similarity remained considerably lower than human–human similarity across instruments and factor-number choices.Mean matched factor congruence used repeated within-human sampling as the baseline distribution.
- The latent-structure differences were observed in both qualitative factor graphs and quantitative factor-congruence analyses.
- The findings indicate that benchmark success may reflect producing correct answers under an evaluation format without establishing the same underlying abilities are measured.
- Using human-designed assessments for LLM evaluation risks invalid score interpretation when validity is assumed to transfer without re-establishment.
- Benchmark design must specify and validate constructs across changing model architectures, not focus only on task difficulty.The discussion also identifies implications for educational applications involving LLM tutors, student simulation, and assessment development.
6 Conclusions & Future Work
The study concludes that these assessment instruments show evidence of measuring different constructs for humans and LLMs. It therefore raises validity concerns about transferring human-oriented score interpretations and calls for broader future research.
- Assessment instruments show evidence of measuring different constructs for humans and LLMs.
- Applying human-designed assessments assumes they measure the same skills in LLMs and support inferences about those skills, raising validity concerns.
- Future research should broaden assessment instruments, expand analytical methods, and incorporate human expert judgment when interpreting results.
Limitations
The study’s findings are constrained by limited reproducibility, external validity, LLM sampling and grouping choices, possible data leakage, and coverage of only one aspect of validity.
- Reproducibility: Non-public datasets constrain absolute reproducibility, although the authors released code, prompts, scripts, and assessment materials upon request.The complete codebase, prompts, and analysis scripts were publicly released.
- Scope: External validity is limited because the results draw on few assessment instruments and specific GenAI tools.
- Sampling and model grouping: The LLM dataset contains 120 responses and groups models into one population despite differing, undisclosed architectures.The LLM sample is relatively small compared with the human sample.
- Evaluation materials: Using proprietary LLMs with non-public materials creates potential data-leakage or future-contamination risks that cannot be fully ruled out.The instrument could become accessible to or influence subsequent system versions.
- Validity coverage: The study examines internal structure but not generalization to related tasks designed to measure the same underlying abilities.