Source-linked AI summary

How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles

Shang Wu, Catarina G Belem, Shuyuan Fu, Mark Steyvers, Padhraic Smyth

arXiv:2608.23543v1cs.AI

TL;DR

AI assistance may improve immediate task performance while weakening skill development when it substitutes for independent reasoning. The paper tests this in a controlled three-phase logic-puzzle experiment with varied AI request costs and a Bayesian latent ability model. Lower-cost access increased AI use, AI users performed worse after assistance was removed, and greater independent effort was associated with larger latent-ability gains.

  • Problem

    The paper addresses whether AI assistance that improves short-term performance also undermines longer-term skill development by reducing independent reasoning.

  • Method

    A controlled three-phase logic-puzzle experiment varied AI request costs and used a Bayesian latent ability model to separate initial ability, post-AI ability, and skill change.

  • Results

    Lower-cost assistance induced more requests, AI users performed worse after assistance ended, and greater independent reasoning was associated with larger latent-ability gains.

  • Takeaways & Limitations

    Skill development appears more closely related to preserving independent reasoning than to AI request frequency itself.

  • Takeaways & Limitations

    The study used a simulated AI assistant and examined short-term skill development in a controlled logic-puzzle task.

Abstract

from arXiv · show

While AI assistance can improve human task performance in the short term, it may also undermine the development of skills in the longer term. We examine this tension in a controlled logic-puzzle experiment involving on-demand AI assistance, where participants complete tasks before, during, and after AI is available. By experimentally varying AI request costs, we find that lower-cost assistance induces more frequent AI use. We also find that participants who request AI assistance during the AI-access phase perform worse at the task after assistance is removed, and their subsequent unassisted performance is overestimated when predicted from earlier AI-assisted performance. We use a Bayesian latent ability model to separate initial ability, post-AI ability, and participant-specific skill change, while estimating how independent reasoning during the AI-access phase relates to skill development. The results show that greater independent problem-solving effort is associated with larger gains in latent ability, consistent with the interpretation that skill development is weaker when AI assistance substitutes for independent reasoning.

1 Introduction

The paper examines whether AI assistance that improves short-term task performance also undermines longer-term skill development. A three-phase logic-puzzle study and Bayesian latent ability model relate AI reliance and independent reasoning to later unassisted performance.

  • Motivation: The study tests whether readily available AI assistance promotes or hinders the independent reasoning needed for skill development.The paper frames AI assistance as potentially improving immediate performance while reducing sustained cognitive effort.
  • Study design: Participants completed an initial AI-free assessment, an intermediate phase with optional AI assistance and varied request costs, and a final AI-free assessment.Comparing Phases 1 and 3 permits analysis of subsequent performance after assistance is removed.
  • Findings: Lower-cost AI access led participants to request assistance more frequently, while AI users performed worse after assistance was removed.The introduction reports that lower-cost access increased requests and that requesters had weaker later performance.
  • Findings: Greater independent problem-solving effort was associated with larger gains in latent ability, whereas request frequency was not associated with skill gains after adjustment.The Bayesian model separates initial ability, post-AI ability, and participant-specific skill change.
  • Implication: The findings suggest that AI assistance is most problematic for learning when it displaces the independent reasoning through which skills are built.The paper distinguishes AI use itself from assistance that substitutes for human reasoning.

2 Related Work

Prior research shows that AI assistance can improve immediate performance while weakening later unaided performance, with effects depending on whether AI substitutes for or complements human reasoning. This paper extends that work by examining request behavior and preserved independent effort in logic puzzles.

  • Prior evidence: Pre- and post-assessment studies examine performance before AI exposure, during AI-assisted work, and after assistance is removed.Related studies use this design to assess whether short-run gains persist without AI.
  • Prior evidence: Evidence from mathematical reasoning, reading comprehension, and writing links AI assistance to reduced persistence, cognitive engagement, or later unaided performance.These findings motivate studying how AI assistance affects subsequent independent work.
  • Mechanism: Cognitive offloading can improve immediate performance but may reduce learning when external tools replace reasoning, memory, or problem-solving processes.The mechanism connects reduced internal effort with weaker acquisition of underlying skills.
  • Mechanism: The consequences of AI assistance depend on whether it prescribes actions or directs attention, distinguishing substitution from complementarity.Prior chess research illustrates how the user’s role changes the effect of assistance.
  • Current study: This paper focuses on the frequency and timing of AI requests and the extent to which participants preserve independent problem-solving effort.The study examines these behavioral mechanisms in a controlled logic-puzzle setting.

3 Experiment

The experiment used time-constrained logic puzzles in a randomized three-condition, three-phase design, with AI assistance available on demand only during Phase 2. Measures captured performance, AI requests, and independent effort for studying subsequent skill development.

  • Task: Participants solved ordering puzzles involving six objects and five constraints, with correctness requiring all six objects in their correct positions.Participants could submit up to two attempts per problem.
  • Procedure: Phase 1 measured baseline performance, Phase 2 introduced optional AI assistance under assigned costs, and Phase 3 repeated the assessment without assistance.Phase 1 lasted 8 minutes, Phase 2 lasted 20 minutes, and Phase 3 mirrored Phase 1.
  • Conditions: Participants were randomly assigned to No-AI, Low-cost AI, or High-cost AI conditions in a between-subjects design.AI requests revealed one randomly selected object’s location and incurred condition-specific point deductions.
  • Conditions: The simulated AI was perfectly correct, ensuring that differences across AI-cost conditions reflected request costs rather than output quality.Every request returned the correct location of a randomly selected object.
  • Sample: The final sample included 124 participants: 42 in No-AI, 43 in Low-cost AI, and 39 in High-cost AI.Recruitment began with 150 U.S.-based English-speaking adults, with 26 exclusions.
  • Measures: Performance was measured using accuracy, response time, and reward rate, defined as average accuracy divided by average response time.Reward rate is expressed as correct objects per minute, with higher values indicating better performance.
  • Measures: AI reliance was measured by total assistance requests, while independent effort was measured by solo share: time spent working independently before requesting assistance.Participant-level solo share averaged problem-level solo shares across Phase 2.
  • Analysis: Bayesian analyses used Phase 1 and Phase 3 performance as signals of latent ability and Phase 2 solo share to predict final skill levels.The analysis examined whether AI-assisted performance predicted later unassisted performance differently across AI-use levels.

4 User Study Results

Participants generally improved on the logic-puzzle task, but AI access and use were associated with weaker subsequent unassisted performance. Bayesian analysis indicates that preserving independent reasoning, rather than request frequency alone, was associated with larger latent skill gains.

  • Skill Development: Unassisted participants increased mean reward rate from 2.03 to 3.86 correct objects per minute, a 90.2% increase.This improvement was statistically significant (N=75, t=7.40, p<0.01).
  • Effects of AI-cost condition: Low-cost AI participants made more Phase 2 requests than high-cost participants, 6.67 versus 3.33 requests.The one-sided Welch test reported t=1.48 and p<0.10.
  • Effects of AI-cost condition: Reward-rate gains from Phase 1 to Phase 3 were 1.95 without AI, 1.66 with high-cost AI, and 1.32 with low-cost AI.The reported increase was smallest in the low-cost AI condition.
  • Post-AI Performance: AI users averaged 3.42 correct items per minute in Phase 3, compared with 3.86 for participants who did not request AI.When Phase 3 performance was predicted from earlier reward rate, AI users were overestimated by 0.22 units, whereas nonusers were underestimated by 0.15 units.
  • Bayesian Analysis: The Bayesian latent-ability model treated accuracy and response time as noisy indicators of initial and post-AI ability while modeling participant-specific skill change.The model incorporated initial ability and Phase 2 solo share to separate latent skill development from observed performance.
  • Independent reasoning and skill development: Solo share was positively associated with latent skill change, with α_solo=0.125, 95% CI [0.002, 0.246], and P(α_solo>0)=0.977.A one-minute increase in solo share corresponded to 1.5% higher predicted Phase 3 accuracy and 2.5% lower predicted response time.

5 Discussion and Conclusion

The study finds that AI assistance is associated with weaker short-term skill development when it displaces independent reasoning. The authors argue that AI support should preserve human thinking to support later unassisted performance.

  • In a controlled logic-puzzle experiment, lower-cost AI requests increased AI use, while requesters performed worse after assistance was removed.
  • Skill development was more closely associated with preserved independent reasoning than with AI request frequency itself.
  • The experiment used a simulated assistant and measured short-term, task-specific skill development in a controlled logic-puzzle task.
  • Because the model used participant-level phase aggregates, it could not separately identify residual differences in problem difficulty or problem exposure.
  • Designing AI systems to support rather than replace human thinking may help preserve subsequent unassisted performance.

A.1 Prior specification

The Bayesian latent ability model uses weakly informative priors, with standardized inputs placing coefficients on a common scale.

  • The model uses weakly informative priors for all parameters after standardizing observed performance measures and predictors.
  • Latent initial ability has a standard normal prior, θ_i1 ∼ N(0, 1).
  • Population-level skill-change coefficients α_0, α_θ, and α_solo each receive normal priors, N(0, 1).

A.2 MCMC inference

The model is fit with Stan using adaptive Hamiltonian Monte Carlo and multiple posterior sampling chains, alongside a non-centered parameterization for participant-specific skill changes.

  • The Bayesian latent ability model is fit in Stan through CmdStanPy using the No-U-Turn Sampler, an adaptive Hamiltonian Monte Carlo algorithm.
  • Four chains use 1,000 warmup and 2,000 post-warmup iterations each, yielding 8,000 posterior draws.
  • A non-centered parameterization separates standard-normal skill-change residuals from their residual scale to improve hierarchical-model sampling efficiency.

A.3 Model Diagnostics and Predictive Checks

The diagnostics indicate satisfactory posterior convergence and mixing, while posterior predictive checks reproduce central tendencies and most observed dispersion. Held-out predictive performance is better for the solo-share models than for the constant-change baseline.

  • All reported R-hat values are below 1.01, effective sample sizes are satisfactory, and no divergent transitions or sampling problems are observed.
  • Observed means for all four Phase 1 and Phase 3 measures fall within 95% posterior predictive intervals, while three of four standard deviations do so.
  • The model somewhat overpredicts dispersion in Phase 1 log response time but reproduces central tendencies and most observed dispersion well.
  • Mean held-out log predictive density is −2.661 (0.090) for the main solo-share model versus −2.734 (0.088) for the constant-change baseline.
  • The solo-share model with cost indicators and the AI-usage model have mean held-out log predictive densities of −2.673 (0.090) and −2.685 (0.090), respectively.

A.4 Alternative specification with solo share and AI-cost condition controls

Adding AI-cost condition controls to the solo-share Bayesian latent ability model does not provide clear additional explanatory information about latent skill change.

  • The alternative model adds randomized low-cost and high-cost AI conditions as direct predictors alongside initial ability and solo share.The no-AI condition is the reference group.
  • Phase 1 and Phase 3 accuracy and log response time remain noisy measurements of latent ability, with post-AI ability defined as θ_i3 = θ_i1 + δ_i.
  • −0.084, 95% CI [−0.399, 0.225] for the low-cost condition indicates a negative but highly uncertain coefficient.The credible interval contains zero.
  • 0.062, 95% CI [−0.238, 0.359] for the high-cost condition indicates a positive but highly uncertain coefficient.The credible interval contains zero.

A.5 Alternative specification with AI usage

Replacing solo share with total AI usage tests whether request frequency explains latent skill change. The posterior estimates indicate that it does not, supporting independent reasoning preserved during Phase 2 as the more relevant behavioral signal.

  • The AI-usage specification replaces solo share with the standardized total number of AI assistance requests during Phase 2.
  • 0.0004, 95% CI [−0.122, 0.124], is the AI-usage coefficient, with P(α_usage > 0) = 0.501.The coefficient is essentially zero.
  • AI usage does not explain latent skill change after accounting for initial ability.
  • −0.495, 95% CI [−0.725, −0.268] for initial ability indicates that participants with lower initial latent ability tended to show larger skill gains.
  • The combined specifications support preserved independent reasoning during Phase 2, rather than request frequency, as the relevant behavioral signal for latent skill development.

A.6 Held-out model evaluation

The study compares four Bayesian latent ability specifications using held-out prediction of Phase 3 outcomes. The main solo-share model provides the strongest overall predictive fit.

  • The four specifications are the main solo-share model, solo share with AI-cost controls, AI usage, and a constant-change baseline.The baseline sets δ_i = α_0 for all participants.
  • The evaluation uses 5-fold held-out prediction, fitting each model on four-fifths of participants and predicting the remaining one-fifth.
  • Held-out log likelihood measures out-of-sample predictive accuracy from posterior predictive probabilities, with higher values indicating better predictive fit.
  • Held-out mean squared error compares posterior mean predictions with observed Phase 3 accuracy and log response time, with lower values indicating better point prediction.
  • The main solo-share model achieves the highest held-out log likelihood and lowest held-out MSE for both Phase 3 accuracy and log response time.These results favor the solo-share model on overall predictive fit.

B Experiment Details

After the three-phase study, participants completed a post-study survey. The analyses use AI-condition responses about reported assistance use to compare them with actual usage logs and identify inconsistent behavior for filtering.

  • The post-study survey was retained from the pilot study for consistency after participants completed the three-phase experiment.
  • Participants rated boredom, confidence, effort, strategy development, learning from correct solutions, and prior experience on 5-point Likert scales.
  • Free-response questions asked about hidden rules, solving strategy, overall experience, and suggestions for improvement.
  • Participants in AI conditions additionally rated assistant accuracy, helpfulness, and when they typically used the assistant.
Loading 2608.23543v1…