Source-linked AI summary
Beware of Metacognitive Laziness: Effects of Generative Artificial Intelligence on Learning Motivation, Processes, and Performance
Yizhou Fan, Luzhen Tang, Huixiao Le, Kejie Shen, Shufang Tan, Yueying Zhao, Yuan Shen, Xinyu Li, Dragan Gašević
TL;DR
Hybrid intelligence learning lacks strong empirical evidence about how different agents affect learners’ motivation, self-regulation, and performance. This randomized study compared 117 university students supported by ChatGPT, a human expert, writing analytics tools, or no extra tool. ChatGPT improved essay scores but not intrinsic motivation, knowledge gain, or transfer, while different supports produced different self-regulated learning processes and may encourage metacognitive laziness.
Problem
Hybrid intelligence remains poorly understood, particularly regarding the mechanisms and outcomes of hybrid human-AI learning based on strong empirical research.
Method
A randomized laboratory study compared learners’ motivation, self-regulated learning processes, and performance across ChatGPT, human-expert, writing-analytics, and no-support groups.
Results
ChatGPT significantly improved essay scores over other groups, including human-expert support, but did not significantly improve knowledge gain, transfer, or intrinsic motivation.
Takeaways & Limitations
Different supports can produce different self-regulated learning processes and performance, while ChatGPT may promote technology dependence and metacognitive laziness.
Takeaways & Limitations
The study’s task may not capture the diversity of cognitive and metacognitive processes involved in varied learning activities.
Abstract
from arXiv · showhide
With the continuous development of technological and educational innovation, learners nowadays can obtain a variety of support from agents such as teachers, peers, education technologies, and recently, generative artificial intelligence such as ChatGPT. The concept of hybrid intelligence is still at a nascent stage, and how learners can benefit from a symbiotic relationship with various agents such as AI, human experts and intelligent learning systems is still unknown. The emerging concept of hybrid intelligence also lacks deep insights and understanding of the mechanisms and consequences of hybrid human-AI learning based on strong empirical research. In order to address this gap, we conducted a randomised experimental study and compared learners' motivations, self-regulated learning processes and learning performances on a writing task among different groups who had support from different agents (ChatGPT, human expert, writing analytics tools, and no extra tool). A total of 117 university students were recruited, and their multi-channel learning, performance and motivation data were collected and analysed. The results revealed that: learners who received different learning support showed no difference in post-task intrinsic motivation; there were significant differences in the frequency and sequences of the self-regulated learning processes among groups; ChatGPT group outperformed in the essay score improvement but their knowledge gain and transfer were not significantly different. Our research found that in the absence of differences in motivation, learners with different supports still exhibited different self-regulated learning processes, ultimately leading to differentiated performance. What is particularly noteworthy is that AI technologies such as ChatGPT may promote learners' dependence on technology and potentially trigger metacognitive laziness.
Practitioner Notes
The paper reports that learning support can alter self-regulated learning processes and short-term task performance without changing post-task intrinsic motivation. It also highlights possible technology dependence and recommends active metacognitive engagement when using AI.
- AI technologies such as ChatGPT may promote technology dependence and potentially trigger metacognitive laziness.
- ChatGPT significantly improved short-term essay performance but did not significantly improve intrinsic motivation, knowledge gain, or transfer.
- Learners using AI should deepen understanding and actively evaluate, monitor, and orient their learning rather than follow ChatGPT feedback blindly.
- Teachers should select suitable AI-supported tasks, stimulate intrinsic motivation, and develop scaffolding for active learning.
- Future research should use multi-task and cross-context studies to examine how learners can ethically and effectively learn, regulate, collaborate, and evolve with AI.
1 | INTRODUCTION
Hybrid intelligence research examines how humans and AI can work together in learning, but empirical understanding of these interactions remains limited. This study compares multiple learning agents to identify differences in motivation, self-regulated processes, and performance, while considering risks of cognitive offloading and metacognitive laziness.
- Hybrid intelligence research remains nascent, with limited empirical understanding of mechanisms and outcomes in hybrid human-AI learning.
- Self-regulated learning involves forethought, performance, and self-reflection, while metacognition includes strategies such as goal setting, monitoring, and evaluation.
- Learners face regulatory challenges related to inadequate metacognitive strategies, low achievement motivation, and task complexity, making external support important.
- Prior GenAI research reports potential learning benefits, but also limitations involving contextualised explanation, teacher replacement, hallucination, skill atrophy, and over-reliance.
- Cognitive offloading may reduce internal cognitive engagement and self-regulation; avoiding difficulty may reinforce less effortful decisions and metacognitive laziness.
- The study conducted a randomized laboratory comparison of four support conditions: an AI chatbot, human expert, writing analytics tools, and no support.
- The study’s originality lies in comprehensively comparing motivation, self-regulated learning processes, and performance across the four groups.
2 | BACKGROUND
The background reviews mixed evidence on AI and motivation, emerging evidence that agents shape self-regulated learning processes, and a gap in direct comparisons of AI, human experts, and other tools. The study therefore examines motivation and learning processes across instructional-agent conditions.
- Existing research on AI, human tutors, and learning tools has produced mixed and inconclusive findings about learning motivation.
- The study addresses the need for comparative evidence on AI, human tutors, and checklist tools’ effects on intrinsic motivation.
- The first research question asks whether and to what extent varied learning-support agents influence learners’ intrinsic motivation toward the task.
- Prior studies suggest AI can influence learning behaviours, engagement, and behaviour sequences, including more systematic problem-solving with AI assistance.
- Process-mining and epistemic-network approaches can model and visualise frequencies and transitions between self-regulated learning processes.
- Research has insufficiently compared how learners’ behaviours and self-regulated strategies differ with AI, human tutors, or other tools in the same context.
3 | METHODS
The study randomly assigned 117 university students to four conditions during a two-stage English reading-and-writing task, then compared motivation, self-regulated learning processes, and performance. Support differed only during revision: ChatGPT, a human expert, checklist tools, or no additional support.
- Experimental design: 117 university students were randomly assigned to control, ChatGPT, human-expert, or checklist-tool groups.The groups included 30 control, 35 ChatGPT, 25 human-expert, and 27 checklist-tool participants.
- Experimental design: Participants completed a two-stage English reading-and-writing task involving reading materials, essay writing, and revision.The essay addressed the future of education in 2035 and was evaluated using a provided rubric.
- Experimental design: The control group received no additional support, whereas the other groups used ChatGPT 4.0, a human academic-writing expert, or Checklist Tools during revision.Checklist Tools provided feedback on spelling and grammar, academic style, originality, and rhetorical structure.
- Measures: Intrinsic motivation was assessed with the Intrinsic Motivation Inventory and compared across groups using ANOVA followed by Tukey’s HSD.The inventory covered interest/enjoyment, perceived competence, effort/importance, and pressure/tension.
- Measures: Learning traces captured navigation, clicks, mouse movement, and keystrokes, which were parsed into learning actions and self-regulated learning processes.Kruskal-Wallis and Mann–Whitney tests compared process frequencies, while pMineR with a first-order Markov Model examined process sequences.
4 | RESULTS
The four support conditions did not differ significantly in post-task intrinsic motivation, but they produced different self-regulated learning frequencies and sequences during revision. ChatGPT produced the largest essay-score improvement, without significant group differences in knowledge gain.
- Intrinsic motivation: No significant group differences appeared for Interest/Enjoyment, Perceived Competence, Effort/Importance, or Pressure/Tension.The reported statistics were F=1.087, p=0.358, η2=0.029; F=0.453, p=0.716,η2=0.012; F=1.152, p=0.332, η2=0.030; and F=0.546, p=0.652,η2=0.015, respectively.
- Intrinsic motivation: Descriptively, the checklist group reported the highest interest, enjoyment, perceived competence, and effort, with the lowest pressure and tension.These descriptive differences were not statistically significant.
- Self-regulated learning processes: During the first stage, self-regulated learning-process frequencies were largely similar across groups, except for slightly lower control-group orientation in the checklist group.The groups received no differentiated support during this stage.
- Self-regulated learning processes: During revision, the ChatGPT, human-expert, and checklist groups showed more Elaboration and Organisation than the control group.These processes primarily involved writing activities.
- Temporal process models: The ChatGPT group frequently looped between interacting with ChatGPT and revising or evaluating essays, whereas the control group connected revision more with reading and task instructions.The human-expert group showed stronger connections between revision, reading, orientation, and evaluation.
- Learning performance: The AI group had significantly greater essay-score improvement than the control, human-expert, and checklist groups.Mean differences were 1.970 versus control, 2.120 versus human expert, and 2.200 versus checklist; the overall score-improvement test was F=4.549, p=0.005, η2=0.108.
- Learning performance: No significant group differences were found in knowledge gain or knowledge transfer.Knowledge-gain comparisons were F=0.913, p=0.438, η2=0.030 for post-test scores and showed no significant differences in knowledge gain; knowledge transfer comparisons were F=0.019, p=0.996,η2=0.000.
5 | DISCUSSIONS
The study found that support conditions shaped self-regulated learning processes and short-term writing performance even when intrinsic motivation did not differ. ChatGPT improved essay scores but was associated with greater reliance on AI and no significant advantage in knowledge gain or transfer.
- RQ1: Intrinsic motivation: Across groups, intrinsic motivation did not differ significantly, yet learning processes and essay-score improvement did.The authors interpret this as evidence that support can affect processes and performance without significant motivational differences.
- RQ2: SRL processes: The checklist tools significantly increased evaluation processes, which the authors attribute to their diagnostic design and rubric-guided feedback.The tools were designed to guide learners in evaluating and revising their own writing.
- RQ2: SRL processes: The ChatGPT group’s revision processes centred on AI interactions and included relatively fewer metacognitive processes than the human-expert and checklist groups.Human teachers triggered more associations among orientation, evaluation, and other metacognitive processes.
- RQ2: SRL processes: The authors define metacognitive laziness as dependence on AI assistance, offloading metacognitive load, and less effective association of responsible metacognitive processes with learning tasks.The interpretation is linked to cognitive offloading and reduced internal engagement.
- Implications: The discussion calls for scaffolding that helps learners divide labour with AI ethically and effectively while developing metacognitive skills.The authors frame this as a requirement for responsible hybrid human-AI learning.
- RQ3: Performance: ChatGPT significantly improved essay scores, but knowledge gain and knowledge transfer did not differ significantly between groups.The authors distinguish improved short-term task performance from unchanged knowledge outcomes.
- Implications: The authors caution that ChatGPT may be most effective for tasks with clear requirements and scoring criteria, while deeper application may require active engagement and critical thinking.They recommend that human-AI interaction supplement rather than replace learner-teacher and learner-system interactions.
6 | LIMITATIONS
The study’s generalizability is constrained by its sample, single reading-and-writing task, short time frame, and limited measurement of metacognitive laziness. The authors call for larger, more diverse, multi-task, and longer-term studies.
- Sample and generalizability: The sample comprised 117 university students, with 70% identifying as female, which may limit representativeness and external validity.The authors recommend larger samples with more balanced gender distributions.
- Task scope: The study relied on a single reading-and-writing task, which may not capture self-regulated learning processes in other activities.The authors recommend multiple task types and cross-context research.
- Time frame: The study’s short-term design limits conclusions about how performance effects translate into lasting knowledge gains and skill development.The authors recommend long-term follow-up assessments.
- Measurement: The study lacked targeted and mature measures for assessing metacognitive laziness.Future research should develop measurement protocols for learners’ potential offloading of cognitive and metacognitive responsibilities.
7 | CONCLUSION
The study found that ChatGPT improved short-term essay performance but did not significantly improve knowledge gain, transfer, or intrinsic motivation. It also highlights potential dependence on AI and metacognitive laziness as concerns for hybrid intelligence in learning.
- ChatGPT significantly improved essay scores compared with the other groups, including those guided by human experts.
- There were no significant differences in knowledge gain or transfer across the groups.
- ChatGPT may improve short-term task performance without boosting intrinsic motivation or long-term learning outcomes.
- The findings raise concerns that learners may become overly reliant on AI, potentially hindering self-regulation and deep engagement in learning.
- The study contributes to hybrid-intelligence research by identifying both the potential and issues of learning with generative AI.
- Future research should deepen understanding of how learners learn, regulate, collaborate, and evolve with AI.
1 Examples of Pre-task and Post-task
The section presents example questions used before and after the task, followed by response instructions and items assessing learners’ experiences, performance perceptions, effort, and motivation.
- Pre-task: Pre-task examples assess knowledge of algorithm improvement, artificial general intelligence, and AI applications in hospitals and healthcare.
- Post-task: Participants were asked to recall the experiment and select the response that best matched their view, with no right or wrong choice.
- Post-task: Responses used a five-point scale ranging from very disagree to very agree.
- Post-task: Post-task items assessed enjoyment, interest, perceived performance, proficiency, effort, task importance, and nervousness.
Control group (CN group)
The CN control group did not receive rewriting support during revision and was instead reminded to use the task instructions and rubric. Its essays were evaluated with a 25-point rubric covering writing quality, content, and task requirements.
- CN learners revised in the same learning environment as stage 1, without the rewriting support provided to the other groups.
- Learners were reminded to focus on the task instructions and rubric when rewriting to obtain a higher essay score.
- The essay rubric assigned a full score of 25 points.
- The rubric assessed basic writing, academic writing, originality, word count, and content criteria including AI in education.
- Content scoring included integration of three topics, scaffolding, differentiation practices, and a future vision for education in 2035.