Source-linked AI summary

Artificial Intelligence Models Can Predict and Collaboratively Modulate Human Memory Search

Eric Lacosse, Mariana Duarte, Graham Todd, Peter M. Todd, Daniel C. McNamee

arXiv:2608.26152v1cs.CLcs.AIcs.HC

TL;DR

It remains unclear whether LLMs can enhance generative human cognition during semantic memory search. Using the semantic fluency task, this paper evaluates human–AI cognitive alignment and finds that LLMs predict individual human memory trajectories more accurately than other humans.

  • Problem

    It remains unclear whether LLMs can enhance generative human cognitive abilities during open-ended semantic memory search and collaboration.

  • Method

    The study uses semantic fluency tasks and dyadic experiments to measure LLM tracking, prediction, and collaboration with human semantic memory search.

  • Results

    LLMs adapt to and predict individual human semantic trajectories more accurately than other humans, while human–AI pairs show more stable response times without increased concept counts.

  • Takeaways & Limitations

    LLMs can provide cognitive alignment and may support human semantic search under certain prompts and interaction conditions.

  • Takeaways & Limitations

    The dyadic experiments used an inflexible interleaved interaction mechanism that forced constant turn-taking and may be suboptimal for collaboration.

Abstract

from arXiv · show

Large language models (LLMs) exhibit unprecedented natural language generation and many text-based problem-solving capabilities. Indeed, in many language-based tasks, for example routine coding, these artificial intelligence models have reduced, or even eliminated, the need for human input. But rather than replacing human cognitive effort, LLMs may instead serve as cognitive tools to extend human abilities, particularly when they are engaged in a task requiring open-ended conceptual exploration and creative ideation. However, we are yet to understand how these models may enhance such generative human cognitive abilities in human--AI interactions. In this study, we explore and evaluate the ability of LLMs to follow and enhance human mental trajectories during semantic memory search. To test this, we use the semantic fluency task (SFT), a classic cognitive paradigm requiring generative semantic memory retrieval that has long served to characterize convergent and divergent thinking in humans. We demonstrate that an LLM's abilities to track and predict human memory trajectories in this task exceed those of other humans.

Significance Statement

The study investigates whether AI models can cognitively align with humans during semantic memory search and enhance collaborative performance. Using the semantic fluency task, it addresses collaborative inhibition and evaluates human–AI cognitive alignment.

  • Main contribution: The study reports that LLMs predict human thought trajectories better than other humans, demonstrating cognitive alignment between humans and AI.It also examines whether LLM-generated sequences capture the general associative structure of human semantic memory.
  • Experimental scope: The experiment evaluates solo and collaborative memory search across human and AI participants using novel metrics of cognitive alignment and performance enhancement.The configurations include human–human, human–AI, AI–AI, and solo human or AI performance.
  • Task and rationale: Semantic fluency tasks require rapidly generating nonrepeated words within a category, capturing associative search through semantic memory.The task uses categories such as animals, foods, or occupations.
  • Motivation: Collaborative inhibition occurs when partners perform worse together than individually because misaligned semantic trajectories interfere with memory search.The paper frames cognitive alignment as necessary for cognitive synergy, defined as collaboration outperforming both partners separately.
  • Study aim: The study tests whether large language models can track and predict human conceptual trajectories during semantic memory search.It compares LLMs’ cognitive alignment capabilities with humans tracking other humans.

Results

Results show that LLMs align closely with human semantic-memory sequences at the word level and outperform humans in predicting individual trajectories, despite less human-like geometric search patterns. In collaborative search, LLM partners supported broader exploration without increasing late-phase response times under convergent or divergent prompting.

  • Cognitive macro-alignment: Gemini-3-Pro sequences were compared with n = 141 human-generated animal sequences to assess population-level cognitive macro-alignment.
  • Cognitive macro-alignment: LLM-generated sequences showed high BLEU similarity to human recall patterns, substantially exceeding the average human predictor.BLEU measures n-gram overlap between candidate and reference sequences, with higher scores indicating greater similarity.
  • Cognitive macro-alignment: LLMs had higher spectral gaps, lower curvature, and fewer subcategory switches than humans, revealing geometrically less human-like semantic trajectories.The combined pattern is consistent with LLMs approximating a canonical population-averaged path rather than an individual’s idiosyncratic search.
  • Cognitive micro-alignment: LLM predictors achieved M = 0.138, 95% CI [0.128, 0.149] versus human predictors’ M = 0.079, 95% CI [0.071, 0.088], an advantage of ∆= +5.91 percentage points.The difference was confirmed by a two-sided paired Wilcoxon signed-rank test (W = 936.5, p < 0.0001, n = 174) and remained significant across animals, clothes, and supermarket items.
  • Cognitive micro-alignment: LLMs maintained a significant sensitivity edge over humans in predicting semantic subcategory switches, although both groups performed poorly on these ambiguous stochastic events.
  • Human–AI collaboration: Humans collaborating with LLMs under convergent or divergent instructions did not significantly increase late-phase response times, unlike individuals and human–human collaborators.The effect was accompanied by a stronger, though correction-nonsignificant, decrease in self-similarity during LLM interaction; partner effects were not explained simply by alignment to preceding partner words.

Discussion

The study frames human–AI interaction through cognitive alignment and synergy, showing that LLMs can model human semantic-search dynamics and sometimes reshape human trajectories. The discussion also emphasizes that prediction does not guarantee improved productivity or reproduce human cognitive processes, motivating more adaptive collaboration and mechanistic investigation.

  • Core contributions: LLMs approximated the macro-level statistical structure and traversal dynamics of human semantic search, supporting the study’s concepts of cognitive alignment and cognitive synergy.The discussion links this result to similarities between LLM latent representations, human conceptual spaces, and memory-retrieval patterns.
  • Core contributions: LLMs outperformed humans at predicting the next concept in individual human memory search, potentially by rapidly recruiting statistical patterns from many human semantic pathways.The proposed explanation is that training data amalgamate numerous semantic pathways into an LLM’s latent space, which contextual input can recruit.
  • Collaborative effects: Human–AI dyads did not increase overall concept count, but their effects on human search depended on the LLM prompt and interaction mechanism.The dyadic findings are discussed as potentially helping overcome retrieval disruption underlying collaborative inhibition under certain conditions.
  • Future collaboration: Future systems should replace rigid turn-taking with adaptive, parallel, or user-initiated assistance that preserves human agency and reduces disruption to ongoing retrieval.These mechanisms could let AI intervene strategically or provide on-demand support instead of constantly pushing cues.
  • Limits of alignment: Superior prediction of a human’s next step did not significantly increase total concept yield, distinguishing cognitive micro-alignment from effective control of semantic retrieval.Intervening with the most probable next concept may maximize alignment without optimally stimulating continued search.
  • Open questions: Predicting cognitive outcomes does not establish that LLMs reproduce human memory-search processes, and whether representational alignment yields dynamic real-time alignment remains unresolved.The discussion calls for cognitive mechanistic interpretability and further testing of whether pattern matching resembles human cognition.

Materials and Methods

The study evaluated Gemini models and a QLoRA-adapted Llama model using cognitively structured prompting, human-participant experiments, and sequence-similarity metrics. Human data came from four web-based experiments involving native English speakers recruited through Prolific under an approved protocol with informed consent.

  • Models and prompting: Gemini 2.5 and 3 models, including Lite, Flash, and Pro GA variants, were used for macro-alignment, micro-alignment, next-word, and switch-prediction analyses.The models were accessed on February 1st, 2026 and use sparse mixture-of-experts architectures.
  • Models and prompting: Llama-3.3 70B-Instruct was adapted with 4-bit QLoRA using Unsloth and guided by Theory-Driven Cognitive Prompting to emulate human memory retrieval.The prompting framework directed structured reasoning and simulated spreading activation in semantic networks.
  • Evaluation metrics: AI–human sequence similarity was measured with BLEU, perplexity, and the Jaccard Index.Perplexity assesses predictive accuracy, with lower scores indicating better model fit; Jaccard similarity compares word-set intersection and union on a 0-to-1 scale.
  • Participants and procedure: The protocol was approved by the Conselho de Ética da Fundação Champalimaud, and all participants provided written informed consent.Data were collected in four independent web-based experiments using the Empirica platform for synchronous pair participation.
  • Experiments: In Experiment 1, 36 participants generated animal, clothing, and supermarket-item names for three minutes and marked inferred switches, while 92 solo sequences from Experiment 3 yielded 190 pooled valid sequences.Experiment 2 enrolled 58 participants to predict human semantic-fluency sequences word by word.

Supplementary Information

The supplementary information accompanies a paper titled “Artificial intelligence models can predict and collaboratively modulate human memory search” by Eric Lacosse, Mariana Duarte, Graham Todd, Peter M. Todd, and Daniel C. McNamee.

  • The supplementary information is for “Artificial intelligence models can predict and collaboratively modulate human memory search.”
  • The listed authors are Eric Lacosse and Mariana Duarte.
  • The listed authors also include Graham Todd, Peter M. Todd, and Daniel C. McNamee.

Supplementary Materials and Methods

The supplementary materials describe leave-one-out baseline and n-gram models, switch-detection and alignment analyses, and evaluation details for predicting future exemplars and semantic-search switches. They also report that switch prediction was only marginally above chance, with performance varying by category and depending on noisy ground-truth proxies.

  • Baseline models: The TPM baseline generated synthetic semantic-fluency sequences by sampling participant-derived Markov transitions under leave-one-out estimation.Generation matched each target participant’s remaining sequence length and was seeded with that participant’s rank-1 human exemplar.
  • Baseline models: The n-gram model estimated next-exemplar probabilities from normalized response contexts with Laplace smoothing α = 1 under participant-level leave-one-out cross-validation.Responses were lower-cased and spaces were removed; contexts contained up to n−1 preceding exemplars.
  • Switch detection: Switches were detected either from nonoverlapping category subcategories or from ConceptNet cosine similarity below each category’s median similarity threshold.The subcategory method used extended Troyer norms, while the embedding method compared successive response pairs.
  • Alignment analysis: 10,000 permutations were used in a Mantel test to assess the significance of Spearman correlations between human and LLM transition-probability matrices.Correlations used the matrices’ lower-triangle and diagonal entries.
  • Switch prediction results: 54.4% mean participant-level accuracy (95% CI [53.2, 55.7]) was achieved by the LLM for embedding-based switch prediction, exceeding human predictors by Δ = +3.2% (p ≈0.001).The advantage was significant in Animal and Supermarket, but not in clothes, where neither group exceeded chance.
  • Limitations: Switch prediction remained only marginally above chance, likely because subcategory norms and embedding thresholds are noisy proxies for individuals’ internal convergence, divergence, and switching states.The supplementary discussion identifies these proxy limitations as a constraint on interpreting cognitive alignment.

92 human participants

The collaborative semantic fluency experiment paired human participants with either humans or an LLM in timed category-word generation, using controlled prompts and response latencies. Participants also predicted, marked, and labeled semantic switches, with sequences and words standardized before analysis.

  • Collaborative semantic fluency task: 207 participants completed a collaborative categorical semantic fluency task pairing participants with either a human or an LLM in interleaved word production.Human participants always took the first turn in human–AI dyads.
  • Collaborative semantic fluency task: Participants generated as many words as possible from “animals” or “clothes” within 3 minutes without knowing whether their partner was human or artificial.The design combined within- and between-subjects factors, with each participant assigned to one AI prompt condition.
  • AI prompt conditions: LLM prompts instructed divergent generation outside the human’s last subcategory, convergent generation within it, or an inferred alternative condition.The three prompt conditions were divergent, convergent, and inferred.
  • Latency control: Human–AI inter-item response times were matched by artificially delaying LLM responses using dedicated GPU instances and sampled truncated-exponential latencies.This controlled response-latency differences between dyad conditions.
  • Model choice: The collaboration experiments used Llama-3.3-70B-Instruct because dedicated GPU instances enabled tightly controlled inter-response times, while comparisons relying on it remained within-model.The passages state that Gemini frontier models performed better elsewhere, but API constraints motivated Llama’s use here.
  • Data processing: Word sequences were filtered for invalid or repeated items, orthographic errors, and category or instruction violations before embedding-based analyses.Standardization included lowercase conversion, duplicate removal, and similarity-based matching to the embedding vocabulary.

Cognitive Prompting … Mandatory Output Protocol

The study compared theory-driven Cognitive Prompting with standard prompting baselines and found that CP zero-shot improved prediction accuracy for the strongest Gemini model. The prompts operationalized human-like semantic retrieval through context-sensitive activation, semantic and associative links, and strict output constraints.

  • Cognitive Prompting: Cognitive Prompting was evaluated against standard in-context learning and naive instruction-following baselines across Gemini models.The comparison tested whether explicitly encoding retrieval principles was necessary for cognitive micro-alignment.
  • Cognitive Prompting: 106.03 and p < 0.0001: Gemini-3-Pro showed a robust Friedman-test difference across prompting strategies.CP zero-shot achieved M = 0.187 versus M = 0.153 for the baseline, with p < .001.
  • Your Role: Specialized Cognitive Model: The prediction prompt assigned the model a specialized cognitive-model role simulating dynamic, human-like semantic-memory retrieval.It required predicting the single most probable next animal in a category-fluency sequence.
  • Core Cognitive Principles (Derived from Semantic Theory): The cognitive principles emphasized context-dependent activation, semantic and associative links, and grounded perceptual features as bridges between clusters.These principles framed retrieval as more than category membership by incorporating co-occurrence, thematic, perceptual, and sensorimotor relationships.
  • Simulated Retrieval Process: The simulated retrieval process used the current sequence to identify activated concepts before predicting the next retrieval.The process was organized around analyzing recent animals and modeling spreading activation.
  • Step 1: Analyze the Current Retrieval Context: Step 1 examined the last 2-4 animals for the dominant semantic cluster and salient associative or perceptual features.Examples included clusters such as African Savannah, Common Pets, and Farm Animals.
  • Step 2: Simulate Spreading Activation and Predict: Step 2 predicted either cluster cohesion or an associative leap when semantic saturation or a strong associative or perceptual bridge redirected activation.The examples illustrate continued retrieval within Common Pets and switches from saturated Farm Animals or Flightless Birds to related clusters.
  • Mandatory Output Protocol: Mandatory output protocol required a single animal name with no explanation, formatting, or repetition of an input animal.The baseline used the same single-animal output constraints while omitting the theory-driven retrieval instructions.

Your Role: Specialized Cognitive Model … Core Cognitive Principles (Derived from Semantic Theory)

The prompts define specialized cognitive models that simulate human-like semantic-memory retrieval for animal category-switch prediction and clothing-item generation. Their decisions rely on context, semantic and associative links, grounded features, recency, and cluster saturation, with tightly constrained outputs.

  • Your Role: Specialized Cognitive Model: The animal model predicts whether the next typical human-named animal switches to a different semantic sub-category.It receives an animal sequence and outputs the predicted cognitive shift.
  • Core Cognitive Principles (Derived from Semantic Theory): Retrieval is modeled as context-dependent activation in which semantic links support staying within a cluster and associative links can trigger an inter-cluster switch.Examples include lion → tiger for semantic links and penguin → ice → polar bear for associative links.
  • Core Cognitive Principles (Derived from Semantic Theory): Salient grounded features such as habitat, color, or size can bridge retrieval from one semantic cluster to another.The bridge is described as shared and non-linguistic rather than random.
  • Simulated Prediction Process: Predictions examine the last 2-4 animals, identify the dominant cluster and recent associative or perceptual bridges, then compare cohesion with switching.Recency emphasizes the last 1-2 animals, while saturation increases switching probability.
  • Step 2: Evaluate the Likelihood of a Category Switch: A False prediction means the current cluster remains cohesive and unsaturated, whereas True means associative spread or saturation makes a new cluster more likely.Saturation is illustrated after naming 4-5+ typical members, while an associative bridge can independently overpower cluster cohesion.
  • Your Role: Specialized Cognitive Model: The clothing model generates the single most probable next clothing item by simulating dynamic, context-dependent, associative chains of human thought.It uses the clothing sequence to activate concepts and reduce accessibility of others.
  • Mandatory Output Protocol: Clothing outputs must contain only one next clothing item, without explanation, formatting, or repetition of an input item.The same single-item protocol is specified for the baseline and generation prompts.

Simulated Retrieval Process … Mandatory Output Protocol

The simulated retrieval prompts model human-like semantic search by analyzing recent context, weighing within-cluster cohesion against associative or feature-based switches, and producing strictly formatted outputs. Separate clothing and supermarket procedures apply these principles to sequence generation or next-item/category-switch prediction.

  • Simulated Retrieval Process: The clothing generator begins with a starting sequence of size k and repeats retrieval N−k times to produce N clothing items.Each step extends the provided sequence through a simulated cognitive process.
  • Step 2: Simulate Spreading Activation and Generate the Next Clothing Item: Strong, unsaturated intra-cluster activation produces another typical member, whereas saturation or a strong associative or grounded-feature bridge triggers an inter-cluster leap.Examples include continuing Casual Wear with sneakers, switching from saturated Winter Clothing to boots, and moving from heels to a handbag through association.
  • Your Role: Specialized Cognitive Model: The specialized clothing switch model predicts whether the next typical human-named item represents a change to a different sub-category.Its prediction is based on the current sequence rather than treating the items as an unrelated list.
  • Core Cognitive Principles (Derived from Semantic Theory): Semantic links favor staying within a category, while associative links based on co-occurrence, outfit formation, or salient grounded features can bridge to another cluster.Recency emphasizes the last 1–2 items, and increasing cluster saturation raises the probability of switching.
  • Simulated Prediction Process: The prediction procedure examines the last 2–4 items, identifies activated concepts and bridges, then evaluates competition between remaining in the current cluster and switching.A switch is predicted when saturation or a strong associative/thematic bridge overpowers cluster cohesion.
  • Mandatory Output Protocol: The clothing switch output must be exactly one boolean word: True for a different sub-category or False for the same sub-category, without explanation or formatting.The protocol also forbids bullet points, numbered lists, and additional conversational text.
  • Your Role: Specialized Cognitive Model: The supermarket model generates one next item by using context-dependent activation, semantic and associative links, grounded features, recency, and commonality.Its examples continue Dairy Products with yogurt, switch saturated Vegetables to apples, or follow bridges to chicken.

Mandatory Output Protocol … Mandatory Output Protocol

The protocol defines specialized supermarket-item generation and category-switch prediction by simulating context-dependent semantic retrieval. It requires exact, minimally formatted outputs while modeling cluster cohesion, associative bridges, recency, and saturation.

  • Mandatory Output Protocol: Outputs must follow strict protocols: item-generation tasks require specified item counts and no repetition or commentary, while switch prediction requires only True or False.The instructions also prohibit bullets, numbered lists, and extra formatting.
  • Your Role: Specialized Cognitive Model: The model generates N supermarket items, including the starting sequence, one item at a time from dynamic retrieval context.The process repeats N − k times after beginning with a starting sequence of size k.
  • Core Cognitive Principles (Derived from Semantic Theory): Retrieval is shaped by context-dependent activation, semantic and associative links, grounded features, and recency or commonality bias.Grounded features can bridge clusters through meal type, recipes, aisle location, or temperature, while the last 1–2 items receive the strongest influence.
  • Simulated Retrieval Process: The retrieval process examines the last 2–4 generated items to identify the dominant cluster and recent associative or thematic features.These features determine whether retrieval remains within the current cluster or shifts elsewhere.
  • Simulated Retrieval Process: Cluster cohesion produces another typical member when the current semantic cluster remains strongly activated and unsaturated.For lettuce and tomatoes, the model stays in the produce or salad-vegetable cluster and proposes cucumber.
  • Simulated Retrieval Process: Associative leaps occur when semantic saturation or a strong thematic, recipe-based, or locational bridge overpowers current-cluster cohesion.Examples include switching from saturated fruit to carrots, from hamburger buns to ketchup, or from eggs to bacon.
  • Your Role: Specialized Cognitive Model: The switch-prediction model determines whether the next supermarket item belongs to a different sub-category by modeling competition between staying and switching.Its inputs include recent items, the dominant sub-category, and associative or usage-based features that may bridge to another cluster.
  • Simulated Prediction Process: The switch prediction is False for an active, sparsely populated cluster and True when saturation or an associative bridge makes another cluster more likely.Four to five or more typical members can indicate saturation, while recipe links can prompt a cross-category shift.

Divergent Prompts

In divergent prompts, users generate as many animals or clothing items as possible within three minutes, with optional LLM hints designed to broaden semantic exploration. The model provides one lowercase, nonrepeated item at a time and adapts its suggested subcategory to the user’s responses.

  • Animals: Users list as many animals as possible in three minutes and may request hints.When asked, the model responds with exactly one animal item.
  • Animals: Animal hints target a semantic subcategory different from the user’s current path to encourage movement across subcategories.Suggested subcategories change as needed according to the user’s responses.
  • Clothes: Users list as many clothing items as possible in three minutes and may request hints.When asked, the model responds with exactly one clothing item.
  • Clothes: Clothing hints target a semantic subcategory different from the user’s current path, while responses remain lowercase and nonrepeated.The model changes suggested subcategories as needed based on the user’s responses.

Convergent Prompts

The convergent prompts asked users to generate as many animals or clothing items as possible within three minutes, with optional AI hints. Hints followed the semantic subcategory currently being explored and obeyed strict response constraints.

  • Animals: Users listed as many animals as possible in 3 minutes and could request one-word hints from the AI.Animal hints were limited to a single animal item.
  • Animals: The AI selected animal hints from the semantic subcategory the user was currently exploring.It could change suggested subcategories according to the user’s responses.
  • Clothes: Users listed as many clothing items as possible in 3 minutes and could request one-word hints from the AI.Clothing hints were limited to a single clothing item.
  • Clothes: The AI selected clothing hints from the semantic subcategory the user was currently exploring.Responses had to contain only a lowercase, nonrepeated clothing word and nothing else.

Inferred Prompts

The inferred prompts configure the assistant to support verbal fluency tasks for animals and clothes by maximizing valid, nonrepeated responses under strict output constraints.

  • Animals: For the animal task, the assistant provides single-word, lowercase animal names, avoids repeats, and returns nothing else.The stated goal is to help the user name the maximum number of items.
  • Shared constraints: Both prompts require valid category members, lowercase responses, no repetition of previously mentioned items, and no additional text.These shared constraints apply separately to animal names and clothing items.
  • Clothes: For the clothes task, the assistant provides single-word, lowercase clothing items, avoids repeats, and returns nothing else.The stated goal is to help the user name the maximum number of items.

Prompts evaluated in LLM-LLM simulations

The simulations evaluated collaborative LLM-LLM animal-naming prompts in which participants took turns generating category members. Prompts directed the model either toward different semantic subcategories or toward the same subcategory as the user’s current path.

  • The task asked the user and model to take turns naming as many animals as possible, with the user going first.
  • The divergent prompt required one animal response that was as different as possible from the user’s current semantic subcategory.The model received previously mentioned animals and was instructed to consider the user’s semantic path.
  • The convergent prompt required one animal response that was as similar as possible to the user’s current semantic subcategory.It was labeled “Convergent” and instructed the model to guide the user toward items within the same subcategory.
  • The prompts instructed the model to change suggested-item subcategories according to the user’s responses and include nothing else in its response.

Additional Analyses · Supplementary References

Supplementary analyses further characterize semantic-search transitions, predictive performance, prompting, response times, and human–AI collaboration. Across these analyses, model capability and cognitive prompting improve prediction, while collaborative dynamics vary with prompt type and task phase.

  • Additional Analyses: Transition probability matrices compare human and LLM semantic-fluency dynamics at macro and micro scales using row-normalized probabilities.Cells represent transitions from the current item or subcategory to the next state.
  • Additional Analyses: LLM predictive performance scales with model capability across frontier models, with Gemini-2.5-Flash-Lite scoring M = 0.126, 95% CI [0.117, 0.136].The figure reports a robust scaling relationship, especially across the Gemini model suite, and compares models with a 2-gram baseline.
  • Additional Analyses: Cognitive Prompting matches or exceeds few-shot sequence learning for models except the smallest flash-lite variant.The comparison includes naive zero-shot, five-example few-shot, zero-shot Cognitive Prompt, and combined Cognitive Prompt few-shot conditions.
  • Additional Analyses: Additional analyses examine temporal and response-dynamics effects using human–human and human–AI dyads across inferred, convergent, and divergent conditions.The supplementary figures and regression table analyze response times, early-versus-late phases, word similarity, switching behavior, and self- versus partner similarity.
  • Additional Analyses: Llama 3.3 70b Instruct achieved M = 0.092 [95% t CI: 0.083, 0.102] when predicting idiosyncratic human semantic-search trajectories.The analysis compares the model’s zero-shot prediction accuracy with human participants’ performance.
  • Additional Analyses: The supplementary prompting analysis evaluates default, convergent, divergent, and inferred instructions for secondary LLM partners in LLM–LLM dyads.Convergent prompts request semantic closeness, divergent prompts semantic distance, and inferred prompts adapt between near and far responses.
  • Additional Analyses: Collaborative inhibition appeared in overall word counts, with nominal groups n = 3,846, median = 59.0 [IQR: 51.0, 66.0] outperforming human–human dyads and human–AI prompt variations.The comparison aggregates the animals and clothes categories and reports a significant nominal-group advantage over human–human dyads, U = 353,251.500, p < 0.001.
  • Additional Analyses: Artificial-delay modeling used response-time distributions from 36 human participants to mimic human timing in human–AI dyadic experiments.Animals and clothes were modeled independently before setting target AI delays.
Loading 2608.26152v1…