Source-linked AI summary

Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task

Nataliya Kosmyna, Eugene Hauptmann, Ye Tong Yuan, Jessica Situ, Xian-Hao Liao, Ashly Vivian Beresnitzky, Iris Braunstein, Pattie Maes

arXiv:2506.08872v2cs.AI

TL;DR

The paper examines how LLM assistance affects cognition during essay writing, a complex educational task requiring multiple mental processes. Using EEG, NLP, questionnaires, interviews, and essay evaluations across LLM, Search Engine, and Brain-only conditions, it reports weaker cognitive engagement and poorer ownership and quotation recall with LLM assistance. The authors argue that AI integration may be better delayed until learners have engaged in sufficient self-driven cognitive effort, while noting limits from the sample and detection precision.

  • Problem

    The study addresses limited evidence about how LLM-assisted essay writing affects brain activity, memory, essay ownership, and writing outcomes in comparison with search-based and unaided writing.

  • Method

    The study compared LLM, Search Engine, and Brain-only essay writing using EEG, NLP analyses, questionnaires, interviews, and human and AI essay assessments across repeated sessions.

  • Results

    LLM-assisted writing was associated with weaker neural engagement, homogeneous linguistic patterns, impaired perceived ownership, and substantially poorer accurate quotation retrieval than search-based or unaided writing.

  • Takeaways & Limitations

    The findings support delaying AI integration until learners have engaged in sufficient self-driven cognitive effort to promote immediate tool efficacy and lasting cognitive autonomy.

  • Takeaways & Limitations

    The study used a limited participant sample from a specific geographical area and several nearby academic institutions, limiting background diversity and gender balance.

Abstract

from arXiv · show

This study explores the neural and behavioral consequences of LLM-assisted essay writing. Participants were divided into three groups: LLM, Search Engine, and Brain-only (no tools). Each completed three sessions under the same condition. In a fourth session, LLM users were reassigned to Brain-only group (LLM-to-Brain), and Brain-only users were reassigned to LLM condition (Brain-to-LLM). A total of 54 participants took part in Sessions 1-3, with 18 completing session 4. We used electroencephalography (EEG) to assess cognitive load during essay writing, and analyzed essays using NLP, as well as scoring essays with the help from human teachers and an AI judge. Across groups, NERs, n-gram patterns, and topic ontology showed within-group homogeneity. EEG revealed significant differences in brain connectivity: Brain-only participants exhibited the strongest, most distributed networks; Search Engine users showed moderate engagement; and LLM users displayed the weakest connectivity. Cognitive activity scaled down in relation to external tool use. In session 4, LLM-to-Brain participants showed reduced alpha and beta connectivity, indicating under-engagement. Brain-to-LLM users exhibited higher memory recall and activation of occipito-parietal and prefrontal areas, similar to Search Engine users. Self-reported ownership of essays was the lowest in the LLM group and the highest in the Brain-only group. LLM users also struggled to accurately quote their own work. While LLMs offer immediate convenience, our findings highlight potential cognitive costs. Over four months, LLM users consistently underperformed at neural, linguistic, and behavioral levels. These results raise concerns about the long-term educational implications of LLM reliance and underscore the need for deeper inquiry into AI's role in learning.

Summary of Results

Across neural, linguistic, and behavioral measures, LLM-assisted writing was associated with weaker cognitive engagement, more homogeneous essays, and poorer ownership and quotation recall than unaided writing. Session 4 further showed contrasting connectivity patterns after switching conditions.

  • Neural results: Brain-to-LLM participants showed higher directed connectivity across all frequency bands than LLM-only Sessions 1, 2, and 3.The reported Session 4 increase included alpha, beta, theta, and delta bands.
  • Neural results: LLM-to-Brain participants showed less coordinated neural effort in most bands after prior LLM exposure.Their essays also showed bias in LLM-specific vocabulary and less distinctive NER and n-gram usage than essays from other sessions.
  • Linguistic results: LLM-assisted essays showed homogeneous ontology, shared n-grams with Search Engine essays, reduced deviation from the SAT prompt, minimal editing, and frequent copy-paste.The LLM group also showed a stronger similarity pattern in embedding analyses.
  • Interpretive context: Prior research described LLM use as reducing cognitive load while potentially compromising deep engagement and the germane load needed to build robust schemas.The study frames this concern alongside evidence that web searching can also burden working memory and discourage internal retention.
  • Behavioral results: 83.3% of LLM-group participants (15/18) failed to provide a correct quotation, compared with 11.1% (2/18) in both Search Engine and Brain-only groups.The LLM group performed significantly worse than both comparison groups, while Search Engine and Brain-only groups did not differ significantly.

Alpha Band Connectivity

Brain-only essay writing engaged stronger and more distributed connectivity than Search Engine writing, especially in alpha, theta, delta, and selected beta networks. These patterns indicate greater internal coordination involving memory, attention, and executive processing.

  • 0.423 vs. 0.288 dDTF values showed stronger significant alpha-band connectivity for Brain-only than Search Engine participants.
  • Brain-only participants showed stronger alpha connections from frontal and temporal sources to temporal and parietal targets, while Search Engine users had few stronger links.
  • 0.417 vs. 0.355 total significant beta dDTF favored Brain-only, although Search Engine participants dominated 11 connections versus 7.
  • 0.644 vs. 0.331 total significant theta dDTF favored Brain-only, with 22 stronger connections compared with 4 for Search Engine participants.
  • 0.588 vs. 0.264 total significant delta dDTF favored Brain-only, with 21 stronger connections compared with 1 for Search Engine participants.
  • Brain-only networks included widespread delta sources and sinks, including stronger right temporal and frontal inputs than Search Engine networks.

Summary

The study compared EEG connectivity during essay writing with no tools, web search, and LLM assistance, including reassignment after three initial sessions. Connectivity patterns differed by tool and frequency band, while switching conditions produced distinct trajectories of engagement.

  • Overall connectivity: Across frequency bands, Brain-only participants showed more extensive and stronger connectivity than Search Engine participants, particularly in delta and theta.
  • Search Engine group: Search Engine writing showed lower slow-rhythm engagement and greater posterior-to-frontal integration, consistent with externally supported information retrieval and visual-executive processing.
  • LLM versus Search Engine: LLM and Search Engine writing exhibited different connectivity profiles: LLM activity was stronger in selected beta, theta, and high-delta measures, while Search Engine activity emphasized posterior-to-frontal flows.
  • Interpretation: The authors characterize LLM assistance and web search as distinct neurocognitive modes involving externally scaffolded automation versus internally managed curation.
  • Session 4: Only 18 participants attended optional Session 4, when participants were reassigned to the opposite condition from Sessions 1–3.
  • Session progression: Repeated Brain-only writing strengthened connectivity across sessions, while prior LLM use was associated with a more intermediate and reduced high-frequency profile during later unaided writing.

Interpretation

Session 4 revealed contrasting neural and linguistic patterns: Brain-to-LLM participants showed stronger connectivity and integrated AI suggestions, whereas LLM-to-Brain participants displayed reduced engagement and repeated LLM-associated language. Across groups, connectivity and essay characteristics differed with external support, though the n-gram–connectivity relationship was examined in only one topic.

  • Neural connectivity: Brain-to-LLM participants showed higher directed connectivity across alpha, beta, theta, and delta bands than the LLM group's earlier sessions.The increase was described as a network-wide spike after three AI-free essays.
  • Session 4 contrasts: Brain-to-LLM essays incorporated common Brain-only and Search Engine n-grams, while some participants also appeared to adopt LLM-suggested phrasing.The ART and PERFECT topics illustrate these contrasting patterns.
  • Session 4 contrasts: LLM-to-Brain participants showed reduced neural coordination and reused vocabulary and structures associated with prior LLM writing.Examples included repeated n-grams in COURAGE, FORETHOUGHT, and PERFECT.
  • Neural connectivity: The LLM group showed the weakest connectivity, including diminished alpha and theta activity in frontoparietal and prefrontal pathways.The findings were interpreted as weak engagement and limited integration or reflection during tool-driven composition.
  • Neural connectivity: Brain-only participants exhibited the highest connectivity and stronger frontal, parietal, and visual-to-prefrontal coupling associated with internally driven reasoning.Their n-grams included reflective and prosocial phrases such as “true happi” and “benefit other.”
  • Limitations: The proposed relationship between n-gram provenance and brain connectivity remains speculative because the analysis covered only one topic.The authors contrasted stronger intrinsic coupling for abstract or value-oriented phrases with reduced integration for externally aided writing.
  • Linguistic patterns: Brain-only essays varied more across participants, whereas LLM essays were statistically homogeneous within topics.The Search Engine group was described as potentially influenced by promoted or optimized content.

Conflict of Interest

The authors report no conflicts of interest and no external funding for the research.

  • The remaining authors declare no conflicts of interest, while Dr. Kosmyna held a Visiting Researcher position at Google at publication.
  • The research did not receive any external funding.

A: List of clusters for Figure 56.

The figure presents clustered interview insights from Session 4 and Sessions 1–3, with participant comments describing essay choices, writing processes, and tool use.

  • PaCMAP defined clusters of interview insights from Session 4 and Sessions 1, 2, and 3.
  • The figure’s quoted insights are ordered from top to bottom and left to right within each cluster map.
  • Participants discussed prompt choices, essay structures, recalled content, time constraints, and personal relevance.
  • One respondent initially struggled to find examples and used ChatGPT to generate examples and combine outputs into an introduction.

B: Dynamic Direct Transfer Function (dDTF) for Alpha band for participants 36 and 43

This figure describes alpha-band dDTF connectivity for two participants writing on the same topic without tools or LLM assistance, alongside aggregated LLM-group connectivity across Sessions 1–3.

  • The alpha-band dDTF display compares Participants 36 and 43, who wrote on the same topic without tools or LLM assistance.
  • The first two rows show dDTF for all pairs of 32 electrodes, totaling 992 connections.
  • Blue indicates the lowest dDTF value and red the highest; the P-value row shows significant pairs below the 0.05 threshold.
  • Aggregated LLM-group dDTF connectivity is averaged across Sessions 1, 2, and 3, with columns separating connections flowing from and to brain areas.

D: Aggregated dDTF for sessions 1, 2, 3 in Search

The figure organizes aggregated dDTF connectivity across Sessions 1–3 for the Search and Brain-only groups by brain area and connection direction.

  • Search-group dDTF connectivity is aggregated and averaged across Sessions 1, 2, and 3.
  • For both groups, columns separate connections flowing from and to each brain area.
  • The first column reports the total connections per brain area, while adjacent columns sum outgoing or incoming connections.
  • Brain-only-group dDTF connectivity is aggregated and averaged across Sessions 1, 2, and 3.

Top 26 significant dDTF

The table lists significant directed-connectivity comparisons between EEG electrode pairs for the LLM and Brain conditions, with associated values and direction patterns.

  • 0.0268734414 versus 0.0059849266 for Fp2→Pz indicates LLM > Brain.
  • 0.0114481049 versus 0.0647280365 for P7→T8 indicates Brain > LLM.
  • 0.0245060250 versus 0.0067248377 for FC1→Pz indicates LLM > Brain.
  • 0.0268790368 versus 0.0061082132 for PO3→Pz indicates LLM > Brain.
  • 0.0148090106 versus 0.0598439611 for T7→T8 indicates Brain > LLM.
Loading 2506.08872v2…