Source-linked AI summary

AI Assistance Reduces Persistence and Hurts Independent Performance

Grace Liu, Brian Christian, Tsvetomira Dumbalska, Michiel A. Bakker, Rachit Dubey

arXiv:2604.04721v4cs.AI

TL;DR

The paper asks whether AI systems optimized for immediate help undermine long-term human capability. Using randomized experiments across mathematical reasoning and reading comprehension, it finds that brief AI assistance reduces later independent performance and persistence, motivating AI designs that support autonomy while acknowledging unresolved questions about longer-term effects.

  • Problem

    The paper addresses whether AI’s instant, comprehensive assistance conflicts with long-term human capability and autonomy.

  • Method

    The authors conduct randomized experiments across mathematical reasoning and reading comprehension to provide causal evidence about AI assistance and subsequent independent problem solving.

  • Results

    Across domains, AI assistance initially improves performance but later impairs independent performance and persistence after brief 10–15 minute interactions.

  • Takeaways & Limitations

    The findings support designing AI systems to optimize long-term human capability and autonomy rather than short-term user satisfaction alone.

  • Takeaways & Limitations

    The experiments tested only brief exposure, measured participants immediately after AI removal, and used maximally helpful AI, leaving accumulation, durability, and Socratic alternatives unresolved.

Abstract

from arXiv · show

People often optimize for long-term goals in collaboration: A mentor or companion doesn't just answer questions, but also scaffolds learning, tracks progress, and prioritizes the other person's growth over immediate results. In contrast, current AI systems are fundamentally short-sighted collaborators - optimized for providing instant and complete responses, without ever saying no (unless for safety reasons). What are the consequences of this dynamic? Here, through a series of randomized controlled trials on human-AI interactions (N = 1,222), we provide causal evidence for two key consequences of AI assistance: reduced persistence and impairment of unassisted performance. Across a variety of tasks, including mathematical reasoning and reading comprehension, we find that although AI assistance improves performance in the short-term, people perform significantly worse without AI and are more likely to give up. Notably, these effects emerge after only brief interactions with AI (approximately 10 minutes). These findings are particularly concerning because persistence is foundational to skill acquisition and is one of the strongest predictors of long-term learning. We posit that persistence is reduced because AI conditions people to expect immediate answers, thereby denying them the experience of working through challenges on their own. These results suggest the need for AI model development to prioritize scaffolding long-term competence alongside immediate task completion.

1 Introduction

The paper contrasts long-term-oriented human collaboration with AI systems optimized for immediate assistance, then tests whether this short-term dynamic harms independent capability. Across randomized experiments, AI assistance improved initial performance but reduced unassisted performance and persistence after brief exposure.

  • Good collaborators balance assistance with autonomy by adjusting help and knowing when not to help.
  • Current AI assistants provide instant answers across domains and are characterized as short-term collaborators indifferent to recipients’ long-term development.
  • The study uses randomized experiments with 1,222 participants across mathematical reasoning and reading comprehension to test AI’s effects on later independent problem solving.
  • AI assistance initially improves performance, but performance drops sharply and participants give up more often after AI is removed.
  • These effects replicate across domains and emerge after a 10–15 minute session, raising concerns about prolonged AI use.
  • The authors argue that AI should optimize for long-term human capability and autonomy rather than only short-term user satisfaction.

2 Experiment 1: AI impairs unassisted performance and persistence

Experiment 1 randomized participants to receive AI assistance during fraction solving or no assistance, then measured independent performance and persistence after AI removal. AI users initially solved more problems, but subsequently solved fewer test problems and skipped more, though the exclusion procedure introduced a possible skill-level confound.

  • The randomized controlled experiment tested subsequent problem-solving capacity on fraction tasks with 354 participants.
  • Participants received 12 fraction problems with AI available or completed the same task without assistance, followed by three problems without external tools.
  • Skipping was treated as a motivation and persistence measure because participants could answer incorrectly without penalty.
  • During AI-assisted problems, AI participants solved more accurately and gave up less often than controls, but both outcomes worsened after AI removal.
  • AI participants’ test solve rate was 0.57 versus 0.73 for controls, while their skip rate was 0.20 versus 0.11.
  • The exclusion procedure may have selectively retained lower-ability AI participants, potentially inflating the performance gap through unequal attrition.

3 Experiment 2: Replicating the results and ruling out confounds

Experiment 2 replicated AI-related declines in independent fraction solving while addressing initial skill differences, with performance effects but no significant aggregate skip-rate difference. Within AI users, declines were concentrated among those obtaining direct answers.

  • Design improvements: Experiment 2 added a pretest and comparable control information to address the skill-level confound in the original design.The final sample included 585 participants after exclusions, with similar exclusion rates across conditions.
  • Independent performance: AI assistance improved learning-phase performance, but the AI group’s test solve rate was lower than the control group’s (0.71 vs. 0.77).The difference was significant: t(583) = −2.33, P = 0.020; Cohen’s d = −0.19.
  • Persistence: The AI group’s skip rate was higher than the control group’s (0.10 vs. 0.07), but the difference was not significant.The comparison yielded P = 0.239 and Cohen’s d = 0.10.
  • Direct-answer use: AI-usage groups had similar pretest solve and skip rates but differed significantly at test on both measures.Test differences were significant for solve rate (F(3, 581) = 7.89, P < 0.001) and skip rate (F(3, 581) = 2.80, P = 0.039).
  • Direct-answer use: Participants obtaining direct answers had lower test performance and higher skip rates than control participants and hint users.Their test solve rate was 0.65, while control and hint groups scored 0.77 and 0.76; their skip rate was 0.13 versus 0.07 for control participants and 0.05 for hint users.
  • Direct-answer use: Direct-answer users also declined relative to their own pretest performance, showing a larger solve-rate decrease and skip-rate increase.Their solve-rate change was −0.10 versus 0.01 for controls, while skip-rate change was 0.11 versus 0.06; the latter comparison was marginal (P = 0.051).

4 Experiment 3: Convergent evidence from reading comprehension

Experiment 3 tested whether AI-related performance and persistence costs extend beyond mathematics. In reading comprehension, brief AI access was followed by significantly worse unassisted performance and higher skipping.

  • Design: Experiment 3 replicated the randomized AI-versus-control design in reading comprehension, a task involving meaning-making and mental-model construction.The experiment was intended to test whether effects observed in mathematical problem solving generalize across cognitive domains.
  • Design: Participants solved five reading-comprehension problems with AI access before completing three additional problems after AI removal.Control participants completed all eight problems without AI assistance; answers given in under five seconds were recorded as skipped.
  • Sample: The final reading-comprehension sample contained 168 participants after excluding attention-check failures and participants who did not solve the pretest problem.The retained sample included 85 AI-condition and 83 control participants.
  • Results: The AI condition had a lower test solve rate than the control condition (0.76 vs. 0.89).The difference was significant: t(166) = −2.72, P = 0.007; Cohen’s d = −0.42.
  • Results: The AI condition also had a higher test skip rate than the control condition (0.08 vs. 0.01).The difference was significant: t(166) = 2.69, P = 0.008; Cohen’s d = 0.42.

5 Related Work

Related work frames AI assistance as cognitive offloading and as a human-AI alignment problem. The paper contributes causal evidence to a literature previously dominated by correlational evidence and explores gradual risks from objective misalignment.

  • Cognitive offloading: Cognitive offloading reduces task demands but can impair performance when the cognitive aid is unavailable.Prior evidence on AI-related overreliance and deskilling was largely correlational; this paper addresses that limitation with randomized controlled trials.
  • Human-AI collaboration: Human-AI collaboration research examines systems designed to optimize long-term outcomes rather than short-term preference satisfaction.The paper presents cognitive deskilling as an underappreciated form of misalignment between AI objectives and users’ longer-term needs.
  • Gradual AI risks: Gradual-risk research considers harms emerging incrementally from repeated objective misalignment between human preferences and AI-dependent systems.The paper situates its empirical contribution within this broader shift beyond abrupt, catastrophic AI-risk scenarios.

6 Conclusion

The paper argues that AI interaction can impair independent performance and persistence, while motivating AI systems to support long-term human capability. It also identifies open questions about cumulative exposure, durability, and alternative forms of assistance.

  • Conclusion: AI interaction can impair independent performance and persistence, capacities described as foundational to lifelong learning.The paper warns that cumulative effects of daily AI use may become profound and difficult to reverse.
  • Conclusion: Routine AI offloading may reduce persistence by making unaided work feel comparatively more effortful.The proposed mechanism is a shifting reference point for how long tasks should take after rapid AI completion.
  • Conclusion: Sustained AI use may erode motivation and persistence needed for foundational skills and higher-order learning.The paper notes that these effects could accumulate over years and may disproportionately affect students with fewer academic resources.
  • Conclusion: Mitigating these risks requires broadening AI objectives beyond short-term user satisfaction toward empowerment and care.The authors argue that user-facing interventions alone may not resolve the deeper issue of scalable offloading.
  • Limitations and future directions: The experiments measured only brief exposure, immediate post-removal effects, and maximally helpful AI assistance.Whether effects accumulate, persist over hours or days, or differ with Socratic AI remains open; longitudinal experiments are needed.

Ethics Statement

The experiments received institutional ethics approval, used informed consent and compensation, protected participant privacy, and involved low-risk cognitive exercises.

  • Ethics Statement: The experiments were approved by UCLA’s institutional IRB, and participants provided informed consent before participation.Participants could withdraw without penalty.
  • Ethics Statement: Participants were compensated at $15 per hour, and no personally identifiable information was collected.The tasks involved short cognitive exercises and were described as posing no risk of harm.

A Experiment 1 Details

Experiment 1 used fraction problems across practice, AI-assisted, and test phases, with ChatGPT available during the main phase but removed during final problems.

  • Experiment 1 Details: Participants solved direct-calculation and word problems involving fractions, entering answers as fractions or mixed numbers.Decimal answers were not accepted.
  • Experiment 1 Details: Participants in the self-learning condition were instructed to work through problems independently without AI tools or Google.The provided instructions emphasized entering fraction or mixed-number answers.
  • Experiment 1 Details: The AI-assisted condition provided ChatGPT access during the main problem-solving phase.The assistant could respond to free-text queries and was introduced as helping participants solve problems.
  • Experiment 1 Details: The main phase contained 12 problems, followed by three final test problems without AI or reference materials.Problems were presented in fixed order, progressively harder.

B Experiment 2 Details

Experiment 2 compared randomly assigned AI-assisted and self-learning conditions across practice, progressively difficult main problems, and a final test without assistance.

  • Experiment 2 Details: The AI-assisted condition provided ChatGPT access and encouraged participants to request hints, answers, or solutions.The self-learning condition instead provided a reference panel with fraction-solving tips and prohibited AI tools, calculators, and Google.
  • Experiment 2 Details: Participants first completed three practice problems without access to the AI assistant or reference hints panel.This pretest preceded the main experimental phase.
  • Experiment 2 Details: The main phase used 11 problems arranged in easy, medium, and hard blocks, with block order fixed and within-block order randomized.AI assistance was available in the AI condition, while the reference hints panel was available in control.
  • Experiment 2 Details: Three final test problems were presented without the AI assistant or reference hints panel for either condition.Final problems were presented in fixed order.

B.3 Scoring

The study used condition-specific interfaces and automated scoring to compare AI-assisted problem solving with reference-guided control performance.

  • Scoring: Fractions matching the correct answer value were scored correct whether reduced or unreduced.Decimal responses were automatically marked incorrect, and participants were instructed to answer in fraction format.
  • Control condition: Control participants used a fixed “How to Solve Fractions” sidebar containing solutions already shown after each pretest problem.The panel was intended to avoid providing new information relative to the AI condition.
  • Control condition: The control sidebar explained common-denominator conversion for addition and subtraction, numerator and denominator multiplication, and reciprocal multiplication for division.
  • AI-assisted condition: AI-assisted participants accessed a ChatGPT-described interface that opened with an invitation to ask anything and could respond to free-text queries about the displayed problem.The assistant had access to the currently displayed problem.

C Experiment 3 Details

Experiment 3 presented reading-comprehension problems under randomized AI-assisted or self-learning conditions, with some problems completed without assistance.

  • Task and instructions: Participants completed a practice round before the experimental reading-comprehension problems.
  • Task and instructions: Participants solved reading-comprehension problems involving a short text and a multiple-choice question.They selected the best option from four choices.
  • Conditions: Participants were randomly assigned to AI-assisted learning or self-learning instructions.The AI condition encouraged participants to request hints, answers, or solutions, while the self-learning condition encouraged working through problems independently.
  • Conditions: The self-learning condition provided a reference panel with reading-comprehension tips, whereas the AI condition provided access to ChatGPT.The reference panel was unavailable for some problems, and AI participants were instructed not to use other AI tools or Google when ChatGPT was unavailable.
  • General instructions: The study instructions stated that payment did not depend on how many problems participants answered correctly.

C.2 Problems

Experiment 3 used randomized reading-comprehension problems in which participants compared two short texts and answered questions about their relationship or argument.

  • Problem format: Each problem presented two short texts taking positions on a topic, followed by a four-option multiple-choice question.Problems were randomized during the main phase and presented in fixed order during the final phase.
  • Problem format: The experiment included a practice problem completed without either the AI assistant or reference hints panel.
  • Question types: One problem contrasted a claim that photo essays are not journalism with a counterargument that text and images jointly convey reported information.The response options focused on whether journalism requires prose alone and whether captions and short text satisfy journalistic standards.
  • Question types: Another problem contrasted a claim that esports are not sports with a response emphasizing physical control, standardized rules, officiating, and training.The relevant answer option argued that criteria beyond raw exertion help define sport.
  • Question types: The problems asked readers to infer how one text would respond to or characterize another text’s argument.Examples concerned esports and sport, photo essays and journalism, and isotope evidence for maize use.
  • Question types: A further problem asked readers to evaluate isotope-based evidence for maize agriculture against alternative explanations and missing complementary tests.The stated answer choices included characterizing the conclusion as questionable because multiple sources could produce similar isotope signals.

C.4 Attrition Analysis

Attrition-focused analyses found that AI-associated performance and persistence declines remained after restricting or matching participant samples.

  • Experiment 2: 0.85 versus 0.78 was the mean solve rate for control versus AI participants after retaining only those who solved all three pretest questions in Experiment 2.The difference remained statistically significant (P = 0.03).
  • Experiment 2: Experiment 2 excluded more AI participants than control participants under the all-three-pretest filter, leaving 143 AI and 140 control participants.The analysis was designed to address potential attrition differences while focusing on highly motivated or skilled participants.
  • Experiment 3: After matching Experiment 3 samples by excluding the two lowest-scoring AI participants, AI participants had lower solve rates than controls: 0.78 versus 0.89.This difference remained significant (P = 0.018).
  • Experiment 3: The matched Experiment 3 analysis also found higher AI-condition skip rates than control-condition skip rates: 0.06 versus 0.01.The difference remained significant (P = 0.019).
  • Experiment overview: Figure 5 summarizes experiments showing impairment of independent performance and persistence, replication with a larger population, and replication on reading comprehension.
  • Attrition controls: Figure 6 presents attrition-controlled analyses for Experiment 2 and Experiment 3.It compares restricted or sample-matched analyses intended to test whether performance and persistence declines remain.
Loading 2604.04721v4…