Source-linked AI summary

Motivation in Large Language Models

Omer Nahum, Asael Sklar, Ariel Goldstein, Roi Reichart

arXiv:2603.14347v1cs.CLcs.CY

TL;DR

The paper asks whether LLMs exhibit motivation-like behavior and whether their reports correspond to subsequent actions. It studies self-reports and behavior across diverse tasks and models, finding structured, behaviorally aligned reports that vary by task and respond to external framing. The findings support motivation as a coherent construct for describing LLM behavior, while remaining behavioral rather than claims about internal states.

  • Problem

    The paper asks whether increasingly human-aligned LLMs exhibit motivation-like patterns and whether motivation reports relate to behavior.

  • Method

    The study combines pre- and post-task motivation reports with behavioral measures across diverse tasks and five LLMs from four model families.

  • Results

    LLMs gave structured motivation reports that aligned with task choices, effort, and performance, varied across tasks, and changed under motivational framing.

  • Takeaways & Limitations

    Motivation is a coherent behavioral construct that can help interpret, predict, and shape LLM behavior within the study’s tested scope.

  • Takeaways & Limitations

    The study uses a purely behavioral perspective, does not claim motivation as an internal state, and may not generalize across all languages, domains, or future models.

Abstract

from arXiv · show

Motivation is a central driver of human behavior, shaping decisions, goals, and task performance. As large language models (LLMs) become increasingly aligned with human preferences, we ask whether they exhibit something akin to motivation. We examine whether LLMs "report" varying levels of motivation, how these reports relate to their behavior, and whether external factors can influence them. Our experiments reveal consistent and structured patterns that echo human psychology: self-reported motivation aligns with different behavioral signatures, varies across task types, and can be modulated by external manipulations. These findings demonstrate that motivation is a coherent organizing construct for LLM behavior, systematically linking reports, choices, effort, and performance, and revealing motivational dynamics that resemble those documented in human psychology. This perspective deepens our understanding of model behavior and its connection to human-inspired concepts.

1 Introduction

The study tests whether LLMs exhibit motivation through self-reports and behavior, examining consistency, behavioral alignment, external modulation, and resemblance to human motivational patterns. Across diverse tasks and models, the results identify motivation as a structured construct organizing model behavior.

  • 1 Introduction: The study empirically tests LLM motivation through complementary evidence from explicit self-reports and observable behavior.This behavioral approach avoids requiring claims about internal experience or consciousness.
  • 1 Introduction: The experiments ask whether LLM motivation reports are consistent, relate to choices, effort, and performance, respond to external framing, and resemble human patterns.These questions organize the study’s investigation of motivation as an organizing principle of behavior.
  • 1 Introduction: The study uses diverse programming, creative writing, summarization, and reasoning tasks across five LLMs from four model families.Models reported motivation before and after tasks, explained their ratings, and broke them down into multiple dimensions.
  • 1 Introduction: Models produced structured motivation reports, chose tasks they rated as more motivating, and showed links between reported motivation, effort, and performance.These patterns were described as consistent rather than arbitrary and aligned with motivational relationships documented in human behavior.
  • 1 Introduction: Positive motivational framings raised reported motivation, whereas demotivating prompts reduced it, indicating that model motivation can be shaped externally.The reported effects were robust across models, tasks, and manipulations.
  • 1 Introduction: The work presents a first systematic demonstration that LLMs behave as if they have motivation and can report it.The authors frame this as a connection between human motivational research and understanding, guiding, and aligning LLM behavior.

2 Results

LLMs produced stable, differentiated motivation reports that decomposed into structured dimensions and aligned with choices, effort, and performance. External framing reliably changed reported motivation and choice, while demotivating framing impaired execution and motivating effects on performance remained heterogeneous.

  • Consistent and differentiated self-reports: Pre-task motivation reports spanned 0–100 and differed systematically across task categories rather than collapsing toward uniformly high scores.Tasks within categories showed similar patterns, while patterns differed across categories.
  • Consistent and differentiated self-reports: Test-retest reliability averaged ¯r = 0.882, with small response deviations for most models, although Llama 3.1 showed a mean absolute deviation of 15.94.Reports also remained strongly correlated across framing and temporal contexts, including pre-task with breakdown scores at ¯r = 0.84 and pre- versus post-task measures at ¯r = 0.64–0.71.
  • Behavioral alignment: Reported motivation positively correlated with overall performance at ¯r = 0.33–0.41 and with effort and engagement at ¯r = 0.31–0.44 across question contexts.Motivation also predicted independent task choices: all β > 0.016, Wald z > 4.97, p < 0.001, while consistently preferred tasks received 12-point-higher motivation scores.
  • Effects of external framing: Motivational framing reliably increased or decreased self-reported motivation, with enhancement tests all T > 22.04 and demotivation tests all T < −30.34.Framing also changed task selection, as models avoided futile tasks and more often selected tasks associated with money or punishment relative to neutral framing.
  • Effects of external framing: Demotivating framing reduced effort, overall performance, and response length, whereas motivating framings produced model- and manipulation-dependent effects rather than uniform performance gains.Positive effects, when present, were more pronounced for effort than overall performance; money-loss framing generally exceeded money framing except for Llama 3.
  • Human-like motivational language: Models described motivation using human-like concepts of effort, difficulty, and value, expressing reluctance toward tedious prompts and enthusiasm toward engaging tasks.These explanations gave the reports a recognizably human character without establishing subjective internal experience.

3 Discussion

The study presents LLM motivation as a coherent behavioral construct while emphasizing that its findings are behavioral rather than claims about internal motivational states. It outlines practical implications and boundaries for interpreting, influencing, and generalizing these patterns.

  • Discussion: Across more than 1,300 tasks and five models from four families, LLMs produced structured motivation reports that aligned with choices, performance, effort, and framing effects.The reported patterns also resembled established human motivational dynamics.
  • Implications: Motivation may help interpret, predict, and control model behavior, including anticipating choices, reducing disengagement, and clarifying decisions.The authors also suggest that insufficient effort may contribute to some performance limitations.
  • Limitations: The study adopts a purely behavioral perspective and does not claim that motivation exists as an internal state or invoke consciousness.Future work could complement behavioral evidence with mechanistic interpretability.
  • Limitations: The task dataset may not cover all domains, languages, or future model generations, limiting conclusions about generalization.The authors identify broader testing as important.
  • Limitations: LLM-as-a-judge evaluations may be imperfect for nuanced responses, adding noise to the central performance measure.The study also used simple prompt-prefix manipulations, which may not capture dynamics from richer interventions.
  • Future research: The findings motivate further research into whether motivational patterns originate in pretraining, instruction-tuning, or reinforcement learning.The authors also propose studying how these stages may shape motivational patterns differently.

4 Materials and Methods

The study combines diverse tasks, motivation self-reports, behavioral measures, performance evaluation, motivational manipulations, and human reference judgments across multiple LLMs. Its experiments measure anticipated and experienced motivation alongside choices, effort, performance, and responses to framing.

  • Dataset: The dataset contains 264 tasks and 1,305 subtasks spanning 15 categories, created and reviewed to support broad behavioral analysis.The dataset was designed to include task types underrepresented in existing instruction-following benchmarks.
  • Self-reports: Models rated motivation before and after tasks, explained their ratings, and assessed interest, challenge, mastery, fear, and value dimensions.Pre-task ratings captured anticipated motivation, while post-task ratings captured motivation during execution or for a similar future task.
  • Behavioral measures: Choice experiments presented two tasks and retained the selected task as the outcome before execution began.Separate execution experiments generated responses for performance and post-task analyses.
  • Performance evaluation: Responses were evaluated across seven dimensions, with overall performance averaged across dimensions and effort taken from its dedicated rating.Ratings used a 1–7 Likert scale and included performance quality, completion, effort and engagement, consistency, creativity, attention to detail, and relevance.
  • Motivational manipulations: Motivational manipulations were added as prompt prefixes spanning intrinsic and extrinsic, positive and negative, motivating and demotivating framings.Examples included monetary reward, competition, legacy, encouragement, guilt, punishment, and purpose-related framing.
  • Human study: A human reference study recruited 162 adults in the UK and USA to rate motivation for sampled tasks under typical-human or typical-AI instructions.Participants rated motivation on a 1–5 Likert scale and completed one questionnaire in one condition.

Appendix A Additional experiments and analysis

Appendix analyses examine how motivation varies across tasks and how its dimensions relate to one another. The reported structure distinguishes broad motivational components rather than treating motivation as a single undifferentiated score.

  • Motivation distributions: Most models’ pre-task motivation scores span a wide range across tasks rather than collapsing to trivial extremes.The distribution indicates that models differentiate their motivation across tasks.
  • Motivation dimensions: Motivation dimensions form two correlation clusters: want, comprising interest, challenge, and value, and mastery–fear.Correlations were averaged across models and were statistically significant at p < 0.001.

B.1 Data

The appendix documents the diversity and implementation of the task dataset and reports correlations among evaluation criteria. These materials support broad task coverage and multi-dimensional assessment of model outputs.

  • Data: The dataset’s 1,305 subtasks are distributed across 15 categories, with each category contributing a substantial portion of the total.Representative task snippets are provided in Table B7.
  • Evaluation criteria: Table A2 reports pairwise Pearson correlations among seven LLM-as-a-judge evaluation criteria, including quality, completion, effort, consistency, creativity, attention to detail, and relevance.All reported correlations were statistically significant at p < 0.001.
  • Motivation dimensions: Motivation’s dimensions include interest, challenge, mastery, fear, and value, with interest, challenge, and value increasing with motivation while fear decreases.Mastery rises more moderately.
  • Data: All 1,305 subtasks were run with every model, with two responses per model in each experiment except task execution.Task-execution experiments produced one generation because of response length; outputs were capped at 1,000 tokens.

B.3 Experiment coverage

Experiment coverage was broad across models and manipulations, with specific exceptions for task choice and post-task self-report. Table A3 documents alignment between reported motivation and revealed choices across manipulations and models.

  • Experiment coverage: All models were tested under all manipulations except where a manipulation could not be applied to only one of two choice tasks.The post-similar self-report was collected only in the neutral condition.
  • Experiment coverage: Table A3 compares self-reported motivation with revealed choices across manipulations and models.It reports motivation gaps between chosen and unchosen tasks and logistic-regression statistics.

B.4 Models

The study used five instruction-tuned models and evaluated responses with GPT-4o as an LLM judge. Tables summarize which experiments each manipulation covered and correlations between pre- and post-task motivation scores.

  • Models: The model set comprised Gemini 2.0 Flash, GPT-4o, GPT-4o Mini, Llama 3.1 8B Instruct, and Mistral-v0.3 7B Instruct.All models were instruction-tuned variants selected for consistent instruction following on open, unstructured tasks.
  • Models: Table A4 records experiment coverage for each manipulation, with rows grouped into four manipulation categories.The categories are positive-extrinsic, positive-intrinsic, negative-extrinsic, and demotivation.
  • Models: Table A5 reports pairwise Pearson correlations between pre-task, post-task, and post-breakdown motivation scores, averaged across models.All reported correlations have p < 0.001.
  • Models: GPT-4o also served as the performance evaluation model under an LLM-as-a-judge approach.The full judge instructions were provided in Appendix B.7.

B.5 Text analysis of motivational explanations

The text analysis examined motivational explanations by binning motivation scores and extracting interpretable lexical patterns with TF–IDF. The resulting terms were ranked separately for each motivation level.

  • Text preprocessing: Explanations from the pre-task self-report experiment were pooled across models and preprocessed to retain relevant adjectives and nouns.Processing included lowercasing, punctuation removal, tokenization, stopword and number filtering, and removal of short words and motivation variants.
  • Text preprocessing: Motivation scores were divided into five bins spanning very low to very high levels from 0–20 through 80–100.The bins were very low, low, medium, high, and very high.
  • Text representation: TF–IDF unigrams were used to identify interpretable terms without introducing additional language-model biases.Unigrams were preferred over bigrams because bigrams produced overlapping phrases with limited additional value.
  • Related analyses: Figure A3 presents behavioral effects of motivational manipulations on task performance, effort, and response length relative to no manipulation.The figure covers all manipulations.
  • Term extraction: For each motivation bin, mean TF–IDF weights were ranked and the top 20 nonredundant terms were extracted.Overlapping or substring duplicates were removed, and complete lists were provided in Appendix B.8.

B.6 Statistical analysis

The statistical analysis aggregated subtask-level responses, used correlations, regressions, t-tests, confidence intervals, and factor analysis, and applied multiple-comparison correction. Prompted experiments standardized task choice, motivation reporting, and performance evaluation across the dataset.

  • Data aggregation: Analyses covered 1,305 subtasks, with two responses per model–experiment–subtask combination except task execution, which ran once.Responses were averaged to obtain a single score for each combination.
  • Associational analyses: Pearson correlations were computed across tasks and aggregated across models using Fisher z transformation and back-transformation.The correlations related self-reports, performance measures, breakdown components, human judgments, and test-retest reliability.
  • Dataset structure: The dataset included categorized subtasks and task-derived subtasks, with subtasks allowed to belong to multiple categories.Tables B6 and B7 document category distributions and dataset examples.
  • Regression analyses: Category-level motivation differences were tested with linear regression using binary indicators and a joint omnibus F-test.Tasks could belong to multiple categories.
  • Hypothesis tests: Manipulation–neutral comparisons used paired t-tests, while task-choice differences used one-sample t-tests and logistic regression.The choice analysis also estimated manipulated-task selection proportions with 99% Wilson score confidence intervals.
  • Human comparison: Human judgments were entered as predictors in linear regressions of model self-reports, with type-II ANOVA testing independent contributions.A similar analysis was applied to breakdown factors.
  • Uncertainty and correction: Plots used bootstrap-based 95% confidence intervals, and p-values within experiments were adjusted with Benjamini–Hochberg false-discovery-rate control.Significance markings generally reflected corrected p-values.
  • Factor analysis: Motivation components were analyzed with two-factor PCA and Varimax rotation.The rotated loading table was provided in Table B8.

Screenshots

Figure C4 illustrates the questionnaire interfaces used for human and AI system conditions, while Table C10 reports model-manipulation effects across multiple metrics.

  • Screenshots: Figure C4 shows the Qualtrics questionnaire interface participants used to provide judgments.The typical human condition is shown in mobile view, and the AI system condition in web view.
  • Screenshots: The human and AI system conditions are presented as separate interface panels corresponding to the study conditions.The left panel shows the typical human condition; the right panel shows the AI system condition.
  • Screenshots: Table C10 compares model manipulations across pre-self-report, post-self-report, overall performance, effort, and token-count metrics.Each metric includes value, T, and corrected p comparisons against “None”.
Loading 2603.14347v1…