Source-linked AI summary

Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses

Alona Strugatski, Licol Zeinfeld, Jason Cooper, Shelley Rap, Gil Schwarts, Giora Alexandron

arXiv:2608.17810v1cs.CLcs.AIcs.HC

TL;DR

Human-designed assessments assume that LLMs and humans rely on comparable underlying skills. This study combines exploratory factor analysis with blind expert interpretation and finds that LLM-derived factors were largely uninterpretable to experts, unlike human-derived factors.

  • Problem

    Human-designed assessments implicitly assume that LLM and human performance reflects comparable skills and item-to-skill mappings.

  • Method

    The study applies exploratory factor analysis to human and six LLM responses, then has subject-matter experts blindly interpret the resulting factor graphs.

  • Results

    Experts interpreted all human factors, but none of the quantitative-reasoning LLM factors and only two of four chemistry LLM factors.

  • Takeaways & Limitations

    LLM latent structures may differ from human skill organization and remain opaque to human experts, complicating interpretation of constructs measured by human assessments.

  • Takeaways & Limitations

    The study uses few instruments in one language, pools different LLMs, and analyzes responses with a small number of SMEs without inter-rater agreement.

Abstract

from arXiv · show

The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM performance carry the same substantive, human-interpretable meaning as the cognitive constructs governing human learners. Using responses from humans and six LLMs across quantitative reasoning and chemistry assessments, we conducted Exploratory Factor Analysis (EFA) separately for both groups. Subject-Matter Experts (SMEs) then blindly evaluated the resulting factor graphs to ascribe pedagogical meaning to the emerged constructs. SMEs successfully interpreted most of the human-derived factors. Conversely, they could not ascribe meaning to any LLM-derived factors in quantitative reasoning and interpreted only half of the LLM factors in chemistry. By combining data-driven EFA with blind expert interpretation, this framework shows that LLMs frequently operate on statistically opaque mechanisms distinct from human reasoning.

1 Introduction

The study tests whether latent factors underlying LLM performance on human-designed assessments have the same interpretable meaning as factors underlying human learners’ performance. Using EFA and blinded SME evaluation across quantitative reasoning and chemistry, it finds human factors interpretable but LLM factors largely uninterpretable.

  • Motivation: Human-designed assessments implicitly assume that LLM and human performance reflects the same skills and item-to-skill mappings.The paper frames this assumption as a central concern in evaluating LLMs with standardized exams developed for human learners.
  • Method: The study compares latent factors from human and six-LLM responses to chemistry and quantitative reasoning assessments using exploratory factor analysis.Researchers generated loading graphs separately for human respondents and each LLM version, then presented blinded, shuffled results to subject-matter experts.
  • Findings: SMEs interpreted all factors derived from human learners across both assessments.Human factors were evaluated in relation to established theories, domain knowledge, and expected response processes.
  • Findings: 0 LLM-derived quantitative-reasoning factors were interpretable to SMEs, while 2 of 4 chemistry factors were meaningfully described.The contrast suggests that LLM factor structures may not correspond to human-interpretable cognitive constructs for the examined instruments.
  • Contribution: The contribution is a measurement-oriented evaluation framework that examines the substantive meaning of latent response structures beyond accuracy comparisons.The study reports that LLM latent structures differed from human structures and were largely uninterpretable to human experts.

2 Related Work

LLM evaluation commonly reuses human-designed assessments, but contamination, aggregate scoring, and differences between human and LLM response mechanisms limit the validity of direct comparisons. These limitations motivate finer-grained, data-driven analyses of response patterns and latent structures, alongside evaluation frameworks designed for artificial agents.

  • Assessment reuse: Human-designed assessments are widely reused in LLM evaluation because they are available, often use automatically scorable closed-form items, and enable cross-model performance comparisons.This practice extends instruments originally developed to measure human knowledge or reasoning across domains.
  • Evaluation limitations: Public assessments may contaminate evaluation through training-data exposure, making measured performance reflect memorization rather than domain ability.Aggregate accuracy can also hide context-specific variation, while limited item-level results constrain detailed analysis.
  • Human–LLM divergence: LLMs show divergent, aberrant, and psychometrically implausible response patterns compared with human examinees, including narrow response distributions that persist after calibration.These findings indicate that LLM assessment responses may not resemble human learner profiles.
  • Latent-structure evaluation: Because LLMs compute answers through mechanisms different from human cognition, human assessments can undermine construct validity and motivate evaluation frameworks specifically designed for artificial agents.Candidate approaches include multidimensional IRT, latent class analysis, network-based methods, and exploratory factor analysis for discovering hidden response structures.

3 Methodology

The methodology combines separately fitted exploratory factor analyses of human and pooled multimodal-LLM responses with blind subject-matter-expert interpretation. It uses binary-scored assessment data, predetermined factor-retention criteria, and blinded factor graphs to compare latent response structures.

  • Methodological pipeline: The six-step pipeline collected human and LLM responses, converted them into binary matrices, determined factor counts, fitted separate EFAs, and prepared factors for blind SME evaluation.The procedure was designed to identify latent factors underlying response patterns and examine how SMEs characterize them from highly loading items.
  • Data sources: Human data comprised 931 chemistry students and 979 quantitative-reasoning examinees from two multiple-choice assessments containing textual and multimodal content.The chemistry assessment had 22 items; the quantitative-reasoning section had 20 items.
  • LLM data: LLM responses came from six multimodal models spanning Anthropic, OpenAI, and Google, then were pooled into one LLM respondent group rather than analyzed by model.Pooling targeted broad LLM latent structures and increased variation in response patterns.
  • Factor analysis: Identical factor counts were retained within each instrument—four for chemistry and five for quantitative reasoning—before fitting separate human and LLM EFA models.EFA used minimum residual estimation, and factors were not assumed to be independent.
  • Factor analysis: Although both groups used identical factor counts, their item-loading structures diverged; Q5, Q16, and Q19 shared one LLM factor but split across two human factors.The example assigns the LLM items to F8 and the human items to F2 and F5.
  • Blind SME evaluation: Experts in mathematics and chemistry evaluated randomly numbered factors without knowing whether they represented human or LLM responses, judging explanations by within- and outside-factor item consistency.The evaluation asked what highly loading items had in common and measured explanatory quality using supporting and non-supporting items.

4 SME Analysis Results

SMEs could not satisfactorily explain any LLM-derived factor in quantitative reasoning, while they explained three human factors. In chemistry, factor interpretation relied mainly on structural similarities among items rather than substantive content.

  • Quantitative reasoning: None of the quantitative-reasoning LLM factors could be satisfactorily explained by experts, whereas three human factors were explained.The passage illustrates the human interpretations with Factor F10, identified as reading data from a graph.
  • Quantitative reasoning: Factor F10 was interpreted as an epistemic construct involving reading data from a graph.Items Q9–Q12 all referred to graphically presented data and loaded on this factor, with Q9, Q10, and Q12 strongly and Q11 weakly.
  • Chemistry: Chemistry factor structure could not be adequately explained from content alone after an initial attempt to identify commonalities.The analysis therefore considered structural characteristics of items rather than only their chemistry-related content.
  • Chemistry: Several chemistry factors shared item-structural characteristics rather than substantive content, although one exception had an uncertain basis.It remained unclear whether that factor arose from the items’ nature or their shared content.

5 Discussion

The discussion argues that LLM-derived factors often reflect skills or statistical processes that SMEs cannot meaningfully interpret, challenging what LLM assessments measure. It also identifies limitations involving reproducibility, sample size, model pooling, and expert-rater procedures.

  • Interpretation of LLM factors: SMEs’ inability to interpret most LLM factors suggests that LLMs organize domain knowledge differently from humans.The authors link this misalignment to uncertainty about which constructs LLM assessments measure.
  • Interpretation of LLM factors: Different latent structures may help explain why LLMs perform well on some benchmarks but poorly on tasks SMEs regard as requiring the same skills.The discussion presents this as an alternative explanation for inconsistent benchmark and task performance.
  • Interpretation of LLM factors: Opaque LLM factors may reflect non-human skills or statistical processes that cannot be explained through latent traits.The paper calls for deeper investigation of how LLMs represent and activate domain knowledge during problem solving.
  • Limitations: Non-public datasets limited reproducibility, although they reduced contamination risk that could have biased the results.The authors considered the tradeoff necessary for the study’s design.
  • Limitations: The LLM sample was relatively small, different models were pooled despite architectural differences, and SME analysis lacked inter-rater agreement procedures.The paper characterizes pooling as accepted but debatable and notes that only a small number of SMEs participated.
Loading 2608.17810v1…