Source-linked AI summary
Large Language Models Align with the Human Brain during Creative Thinking
Mete Ismayilzada, Simone A. Luchini, Abdulkadir Gokce, Badr AlKhamissi, Antoine Bosselut, Antonio Laverghetta, Lonneke van der Plas, Roger E. Beaty
TL;DR
Existing brain–LLM alignment studies have focused mainly on passive, non-creative tasks, leaving alignment during divergent creative thinking less understood. Using fMRI from 170 AUT participants and RSA across LLMs, this study finds stage-dependent alignment relationships with model size, creativity performance, and post-training objectives, while noting correlational and model-family limitations.
Problem
Prior brain–LLM alignment studies have predominantly examined passive or non-creative tasks, leaving alignment during divergent creative thinking insufficiently studied.
Method
The study compares RSA-based alignment between fMRI responses from 170 participants performing AUT and LLM representations across model sizes, brain networks, processing stages, and post-training variants.
Results
Brain alignment is positively associated with model size and divergent-thinking performance during prompt processing, while post-training objectives show selective, opposing alignment patterns for high- and low-creativity responses.
Takeaways & Limitations
Brain alignment may provide a lens for evaluating and developing creative cognition in LLMs beyond behavioral benchmarks, while convergent-only post-training warrants scrutiny for divergent thinking.
Takeaways & Limitations
The findings are correlational because compared model variants differ in pretraining data, optimization, fine-tuning corpus, and generation length, and some patterns do not replicate across model families.
Abstract
from arXiv · showhide
Creative thinking is a fundamental aspect of human cognition, and divergent thinking-the capacity to generate novel and varied ideas-is widely regarded as its core generative engine. Large language models (LLMs) have recently demonstrated impressive performance on divergent thinking tests and prior work has shown that models with higher task performance tend to be more aligned to human brain activity. However, existing brain-LLM alignment studies have focused on passive, non-creative tasks. Here, we explore brain alignment during creative thinking using fMRI data from 170 participants performing the Alternate Uses Task (AUT). We extract representations from LLMs varying in size (270M-72B) and measure alignment to brain responses via Representational Similarity Analysis (RSA), targeting the creativity-related default mode and frontoparietal networks. We find that brain-LLM alignment is positively associated with model size (default mode network only) and with idea originality (both networks), with these relationships clearest early in the creative process. We further find that post-training objectives are associated with functionally selective differences in alignment: a creativity-optimized Llama-3.1-8B-Instruct retains alignment with high-creativity neural responses while lacking the positive low-creativity alignment present in other variants, but a reasoning-trained variant shows the opposite asymmetry, with negative alignment to high-creativity responses. Together, these results suggest that brain alignment offers an informative lens on how post-training relates to the neural geometry of human creative thought.
1 Introduction
Creative thinking, especially divergent idea generation, is a central human capacity that LLMs increasingly demonstrate. This study extends brain–LLM alignment research to active creative thinking, examining whether model representations align with neural activity during the AUT.
- Divergent thinking generates varied novel ideas from a single starting point and complements convergent thinking’s focus on one defined solution.
- LLMs perform strongly on creativity benchmarks such as the AUT, sometimes matching or exceeding average human performance.
- Prior brain–LLM alignment research has largely examined passive language processing, simple word association, or abstract reasoning rather than active divergent thinking.
- The study uses fMRI data from participants performing the creative AUT and matched non-creative OCT while comparing their neural representations with LLM representations.
- The study reports that post-training objectives are associated with selective alignment differences, with opposite asymmetries for creativity-optimized and reasoning-trained models.
2 Related Work
Research has established meaningful correspondence between LLM representations and neural activity across several cognitive domains, while newer work shows strong LLM performance on divergent creativity tasks. These developments motivate studying alignment during creative cognition rather than passive comprehension alone.
- LLM representations have been aligned with neural activity during vision, audition, language processing, and computer-program comprehension.
- Frontier LLMs can generate highly original AUT responses that sometimes match or exceed average human performance.
- Creative prompting and fine-tuning or preference-optimization strategies have been proposed to improve LLM creative thinking and problem-solving.
3 Methodology
The methodology matches human and LLM task inputs, extracts representations before and after generation, and evaluates their similarity to fMRI activity with RSA. Analyses distinguish creative and control tasks, brain networks, pooling strategies, and creativity levels.
- Human data: 170 healthy subjects performed both the AUT and matched OCT across 46 stimuli, with responses and fMRI recordings collected for each stimulus.
- Human data: The OCT asks for typical object characteristics, providing a matched non-creative control that shares the AUT’s stimulus, structure, and language.
- Model representations: LLMs receive the same instructions and stimuli as participants, with prompt content fixed in primary analyses and representations extracted at prompt and post-generation stages.
- Analysis factors: Analyses examine the DMN and FPN, use SOM as a control, and partition AUT responses into high- and low-creativity groups using human ratings.
- Alignment analysis: Brain–LLM alignment is computed with Representational Similarity Analysis by comparing human and model representational dissimilarity structures.
- Alignment analysis: Alignment scores are reported as subject medians after normalization by subject-specific fMRI noise ceilings.
4 Experimental Setup
The experiments span open-source instruction-following models from 270M to 72B parameters and compare base, creativity-optimized, supervised, preference-optimized, mathematical, and reasoning-trained variants. Additional analyses vary pooling, generation, prompts, and performance measures.
- Models: The model set ranges from GEMMA-3-270M-IT through QWEN2.5-72B-INSTRUCT and includes Llama, Olmo, Falcon, DeepSeek, and Mistral variants.
- Post-training comparisons: Post-training comparisons include base and instruction-tuned Llama and Qwen models, creativity-optimized Llama variants, and reasoning-trained DeepSeek distillations.
- Post-training comparisons: CRPO supervised-finetuning and direct-preference-optimization ablations share the creativity model’s base model and training data to isolate post-training factors.
- Evaluation and representations: Model performance is evaluated with GEMINI-3-FLASH scoring adapted AUT outputs, while representations use either last-token or mean-token pooling.
- Generation: Response-stage activations cover full reasoning-like outputs, including explicit <think>...</think> traces in DeepSeek-R1-Distill variants, with generation capped at 1024 tokens.
- Controls: Empty-prompt and token-length-matched placeholder-prompt conditions test whether observed effects depend on the task instruction.
5 Results
Brain–LLM alignment is strongest and most systematically related to model properties during early prompt processing, while later response-stage and representation-extraction effects are more conditional. Post-training objectives also produce selective alignment patterns for high- versus low-creativity neural responses.
- Early and response-stage alignment: r = 0.55, p < 0.05 for model size and r = 0.52, p < 0.05 for AUT score in DMN alignment during early cue processing.The size relationship uses last-token pooling, whereas the AUT-score relationship uses mean-token pooling.
- Early and response-stage alignment: Response-stage alignment is more variable, and correlations with model size and task performance weaken after models generate responses.Within the Gemma-3 family, these correlations remain positive at the response stage, suggesting family heterogeneity partly contributes to the pooled weakening.
- Instruction dependence: Task-instruction removal eliminates significant alignment correlations with model size or task performance across stages and pooling strategies.The result indicates that alignment in creativity-relevant networks depends on receiving the human-matched task instructions rather than generic input encoding.
- Layer and pooling effects: Peak alignment depth depends on token pooling: last-token pooling emphasizes early layers, whereas mean-token pooling yields modes in final and early layers.No reliable relationship links peak depth to alignment magnitude, and the authors characterize peak depth as contingent on representation extraction.
- Layer and pooling effects: Layer selection by maximum alignment produces upper bounds rather than unbiased estimates, but cross-validated correction shows the size correlation survives removal of this optimism bias.The optimism bias does not scale with layer count.
- Post-training objectives: During response generation, CRPO-Llama retains positive alignment with high-creativity responses but lacks reliable positive low-creativity alignment, while DeepSeek reasoning distillation shows the opposite asymmetry.The CrPO pattern is absent in matched SFT and DPO variants, and reasoning-training effects partially generalize beyond the Llama family.
- Post-training objectives: QWEN2.5-7B-INSTRUCT shows no consistent lean toward either high- or low-creativity response population.
6 Discussion
The results position brain alignment as a complementary lens on creative cognition in language models and raise questions about post-training strategies centered on convergent tasks. They also emphasize that alignment patterns depend on processing stage and training objective.
- Post-training has emphasized convergent-thinking tasks such as mathematics and coding because they provide readily evaluable correct answers or judgeable outputs.
- Brain alignment may evaluate and develop creative cognition beyond behavioral benchmarks by capturing aspects closer to human creative thought’s computational principles.
- Exclusive post-training on convergent-thinking tasks warrants closer scrutiny for its effects on divergent thinking.
Conclusion
This study examines brain–LLM alignment during divergent creative thinking using fMRI and the Alternative Uses Test. Alignment relates positively to model size and creative performance early in processing, while response generation and post-training objectives produce nuanced, selective differences.
- The study presents the first investigation of brain–LLM alignment during divergent creative thinking, using fMRI from participants performing the Alternative Uses Test and a non-creative control task.
- Alignment with creativity-relevant brain networks is positively associated with model size and creative task performance during prompt processing, but weakens after response generation.
- Creativity-optimized training retains alignment with high-creativity neural responses while lacking positive low-creativity alignment, whereas reasoning-chain training shows negative alignment with high-creativity responses.
- The correspondence between LLM representations and human creative thought is sensitive to processing stage and training objective.
A Limitations
The study identifies important scope and attribution limits: post-training effects are correlational, mechanistic explanations are absent, and instruction tuning and reasoning budget remain confounded.
- Correlational rather than causal evidence: Post-training associations are not causal because model variants differ in pretraining data, optimization, fine-tuning corpus, and generation length.The negative high-creativity alignment found for Llama reasoning distillation does not replicate fully in its Qwen counterpart.
- No mechanistic account of the observed effects: The analyses establish differential alignment patterns but do not explain the mechanisms producing them.Layer-wise, probing, and component-level analyses are identified as needed next steps.
- Instruction tuning and reasoning budget are not isolated: Instruction tuning and reasoning budget are not isolated, preventing a matched comparison of pre- versus post-instruction tuning or thinking versus non-thinking models.The generation setup caps outputs at 1024 tokens but does not match reasoning budgets.
- Scope of the post-training comparison: The post-training comparison excludes a model fine-tuned for domain-general cognitive modelling.Such a model would provide a reference against creativity- and reasoning-optimized objectives.
- Measurement boundary: RSA alignment compares LLM and fMRI representational geometries using signed Spearman correlations between their RDM upper triangles.Alignment scores range from -1 to 1, with negative values indicating anti-alignment.
D Robustness of the Alignment Correlations
Robustness checks support the main DMN scaling association, while AUT-performance correlations are more fragile and control analyses restrict the effects to the intended creative condition.
- Primary correlations: n = 24 models; the DMN–size prompt-stage correlation remains robust across BCa intervals, leave-one-out analysis, and Spearman correlation.r = +0.55, BCa CI [+0.24, +0.77], LOO range [+0.45, +0.61], and Spearman ρ = +0.65 (p < 0.001).
- Leverage point in the AUT-score correlations: Removing GEMMA-3-270M-IT reduces AUT-score correlations from r = +0.52 to +0.40 in the DMN and from +0.46 to +0.29 in the FPN.The positive direction is preserved, but the smallest model is influential for both estimates.
- Multiple comparisons: After Bonferroni correction, DMN–size (pbonf = 0.016) and DMN–AUT-score (pbonf = 0.036) remain significant, whereas FPN–AUT-score does not (pbonf = 0.079).The correction is applied to the pre-specified family of three primary correlations.
- Control conditions: In the matched OCT control, DMN–size correlations are nonsignificant (r = +0.09, p = 0.66; r = +0.22, p = 0.30), unlike AUT (r = +0.55, p = 0.005).No FPN–OCT cell reaches significance, and AUT FPN–size correlations are null.
- Within-family scaling: Within Gemma-3, prompt-stage DMN correlations are r = 0.90 (p = 0.04) for model size and r = 0.85 (p = 0.07) for AUT score.Within Qwen2.5, prompt-stage correlations are r = 0.89 (p = 0.02) for size and r = 0.74 (p = 0.09) for AUT score, but the two largest models influence the size estimate.
- Within-family scaling: Figure 5 reports prompt-only Gemma DMN alignment against model size and task performance using Pearson r and p.The figure concerns the within-family scaling analysis.
F Per-Stimulus and Per-Subject Sample Sizes for the Response Creativity Partitioning
The creativity partition includes responses from every AUT stimulus, but response counts are uneven across stimuli and subjects. Sparse subject-level buckets can make alignment estimates noisy, although median aggregation limits their influence.
- Per-stimulus distribution: All 46 AUT stimuli contributed responses to both low- and high-creativity buckets.High-creativity proportions ranged from roughly 0.66 for comb, candle, and balloon to roughly 0.22 for bucket, dog leash, and brick.
- Per-subject distribution: 2,147 low-creativity responses versus 1,481 high-creativity responses produced an overall imbalance at the 2.0 cutoff.Because each subject completed a fixed number of trials, subject counts form an anti-diagonal distribution.
- Per-subject distribution: Most subjects contributed between 5 and 18 responses to each bucket, but several had fewer than three high-creativity responses and at least one had none.RDMs for sparse buckets are estimated from very few stimuli and are correspondingly noisy.
- Per-subject distribution: Median aggregation limits the influence of subjects with sparse creativity buckets on reported alignment values.The analysis reports these sparse subjects explicitly rather than excluding them.
G Absolute Alignment Levels Across Conditions
Absolute alignment was reliably higher in the creativity-relevant DMN than in the somatomotor control network during prompt processing, whereas AUT and OCT magnitudes were usually comparable. These paired comparisons complement, rather than replace, the correlational analyses of model properties.
- Creativity-relevant versus control network: +0.15 with last-token pooling and +0.16 with mean-token pooling marked higher DMN than SOM alignment at the AUT prompt stage.Both differences survived correction; the last-token result covered 18 of 23 models, while the mean-token result covered 22 of 23.
- Creativity-relevant versus control network: Neither AUT response-stage comparison between DMN and SOM showed a reliable difference.The reported p-values were 0.94 and 0.34 for the two pooling strategies.
- Creative versus non-creative task: +0.14 was the only AUT-over-OCT difference that survived correction: DMN, prompt stage, mean-token pooling.The difference had CI [+0.09, +0.18], occurred in 23 of 24 models, and had p < 10^-6.
- Creative versus non-creative task: Five of eight AUT-versus-OCT cells were null, including DMN prompt-stage last-token pooling, where models split evenly despite median alignment of 0.44 versus 0.34.Two nominally significant cells pointed in opposite directions and neither survived correction.
- Summary: Absolute DMN alignment exceeded SOM alignment at the AUT prompt stage, while OCT alignment was comparable to AUT in most configurations.These magnitude comparisons are complementary to correlational analyses relating alignment to model properties.
H Cross-Validated Layer Selection
Cross-validated layer selection addresses upward bias from taking the maximum alignment across layers. The size–alignment relationship remains essentially unchanged after cross-validation, indicating that layer selection did not manufacture it.
- Bias concern: Maximum-across-layer alignment is biased upward because models with more candidate layers have higher expected maxima.Layer count is strongly collinear with parameter count (r = 0.94), creating a potential size–alignment confound.
- Cross-validation procedure: Five-fold cross-validation selected each model’s layer on other subjects and evaluated alignment on held-out subjects, repeated 20 times.Selection and evaluation never shared subjects, preventing the maximum operation from inflating evaluated values.
- Cross-validation results: 0.13 was the median reduction from naive to cross-validated alignment, equal to 33% of the naive value across a 0.007–0.267 range.Best-layer scores are therefore upper bounds rather than unbiased estimates.
- Cross-validation results: r = −0.48 (p = 0.018) described the relationship between cross-validation reduction and layer count, opposite to the predicted confounding direction.The reduction instead tracked genuine layer structure: the gap correlated negatively with cross-validated alignment (r = −0.81, p < 0.001).
- Size correlation: r = +0.56 (p = 0.004) was the cross-validated DMN–size correlation versus r = +0.55 (p = 0.006) under the naive estimator.The authors conclude that the positive size relationship is not an artifact of layer selection.