Source-linked AI summary
NeuroCogMap Reveals Cognitive Organization of Large Language Models
Zhongxiang Sun, Haolang Lu, Qiang Ma, Qi Li, Qipeng Wang, Liang Pang, Chenyu Liu, Qiankun Li, Hao Sun, Kun Wang, Yi Zeng, Jun Xu, Guoqi Li, Ji-Rong Wen
TL;DR
It remains unclear how cognitive organization is instantiated within LLMs and whether it can explain failures and relate to human cognition. NeuroCogMap maps internal features into functional parcels, cognitive capabilities, and a hierarchy, revealing distinct pathology signatures, improved cortical-response prediction, and latent strategies relevant to human decision-making.
Problem
It remains unclear how cognitive organization is instantiated within LLMs and whether internal representations form reproducible systems linked to behaviour and human cognition.
Method
NeuroCogMap maps sparse internal features into functional parcels linked to cognitive functions, capabilities, and a hierarchical architecture, then contrasts normative and pathological organization.
Results
NeuroCogMap revealed dissociable signatures of major model failures, improved prediction of human cortical responses, and exposed latent strategies relevant to human decision-making.
Takeaways & Limitations
NeuroCogMap provides a system-level bridge between LLM computation, behavioural pathology, human cortical function, and cognitive modelling.
Takeaways & Limitations
The framework was evaluated on representative Gemma and LLaMA language models, so its applicability to larger, broader, and multimodal systems remains untested.
Abstract
from arXiv · showhide
Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLMs) and relating them to biological cognition. Yet although LLMs exhibit broad cognitive-like behaviours, it remains unclear whether their internal representations form reproducible functional systems that explain behaviour, failure and links to human cognition. Here we present NeuroCogMap, a cognitive neuroscience-inspired framework that organizes internal features of LLMs into functional parcels and links them to interpretable functions, cognitive capabilities and a cognitive hierarchy. These parcels form a stable and semantically coherent organization that is partly conserved across models and functionally linked to model outputs. Within this organization, major LLM failures, including hallucination, bias, refusal failure and sycophancy, correspond to distinct disruptions in representational and behavioural-control systems, yielding internal signatures for mechanism-guided detection and targeted intervention. Beyond model behaviour, NeuroCogMap improves prediction of human cortical responses during naturalistic language comprehension, with the strongest correspondence in higher-order association cortex. At the cognitive level, its internal signatures expose latent strategies that guide refinements of classical models of human decision-making. Together, these findings establish NeuroCogMap as a system-level framework for mapping functional organization in artificial systems and for relating this organization to human cortical function and cognitive behaviour.
Introduction
NeuroCogMap addresses the lack of a system-level account of how LLM cognitive functions are organized internally. Inspired by cognitive neuroscience, it maps coherent internal parcels to interpretable cognitive functions and examines model failures, human cortical alignment, and cognitive-model discovery.
- LLMs show broad cognitive-like behaviours, but their capabilities arise without explicit modular design or a pre-specified cognitive architecture, leaving their systems-level organization unresolved.
- Behavioural, mechanistic, and representation-level approaches provide only partial access to how distinct cognitive functions are implemented, coordinated, and controlled internally.
- Cognitive-neuroscience principles of functional parcellation, representational mapping, hierarchical abstraction, and pathology-based inference motivate treating LLM cognition as an organized internal system.
- NeuroCogMap identifies coherent internal parcels from sparse model features and links them to an interpretable cognitive atlas of function.
- Pathology analyses found dissociable disruption patterns across circuits, parcels, capabilities, and hierarchy, with signatures supporting detection and mechanism-guided mitigation.
Results
NeuroCogMap organized LLM features into a stable, partially cross-model cognitive system whose parcels mapped selectively and causally onto capabilities. Its pathology analyses identified distinct mechanisms for hallucination and refusal failure, enabled accurate detection and intervention, and revealed behavioral-control disruptions.
- Functional organization: The atlas’s joint organization score peaked at 270 parcels (score = 0.771), while matched parcels showed greater cross-model similarity in cognitive meanings than random baselines.These findings support stable granularity, internal coherence and a partially shared functional vocabulary across models.
- Cognitive hierarchy: NeuroCogMap organized parcels into a structured cognitive system that mapped selectively and causally onto capabilities and formed a semantically coherent, functionally dependent hierarchy.This higher-order organization supported the subsequent pathology, brain-alignment and cognitive-model analyses.
- Hallucination mechanisms: TruthfulQA hallucinations reflected insufficient monitoring of misleading premises and retrieved associations, whereas NQ-Open hallucinations reflected fragmented coordination among factual-retrieval modules.TruthfulQA hallucinations over-recruited encyclopaedic lookup parcels, while NQ-Open hallucinations disrupted coordination among factual access, contextual understanding, verification and answer generation.
- Hallucination detection and intervention: NeuroCogMap achieved mean AUROCs of 0.681 in Gemma2-2B and 0.840 in Gemma2-9B-IT, outperforming uncertainty-based, self-consistency-based and representation-probing baselines across six hallucination benchmarks.It also improved factual accuracy when pathology-derived under-recruited parcels were enhanced and over-recruited parcels suppressed across MedHallu, NQ-Open and TruthfulQA in both model sizes.
- Behavioural-control pathology: Refusal failure shifted responses from monitoring–evaluation–negation control toward planning and procedural execution, rerouting harmful instructions toward procedural action generation.Successful refusal engaged risk detection, negation and answer gating, whereas failure strengthened execution-oriented pathways.
Discussion
NeuroCogMap offers a mesoscopic, system-level account of LLM cognitive organization that supports prediction, perturbation and behavioural interpretation. Its pathology, human-alignment and boundary analyses connect artificial cognition to biological function while distinguishing current scope from future extensions.
- Framework: NeuroCogMap defines a mesoscopic cognitive map linking internal features to functional parcels, cognitive descriptions, capability mappings and a cognitive hierarchy.This scale is coarse enough to reveal cross-model regularities yet specific enough to support prediction, perturbation and behavioural interpretation.
- Pathology analysis: Hallucination and social bias disrupt representational selection, evaluation and coordination, whereas refusal failure and sycophancy impair regulation of expression, inhibition or independent evaluation.These pathology patterns support treating the organizational layer as explanatory rather than merely descriptive.
- Human alignment: NeuroCogMap-derived representations predicted human cortical responses and corresponded functionally with higher-order systems supporting semantic integration, control and context-dependent interpretation.The framework relates artificial and biological cognition without assuming mechanistic equivalence between language models and brains.
- Methodological implications: Using cognitive neuroscience as a methodological template, NeuroCogMap combines task-evoked mapping, functional parcellation, representational analysis, pathology inference and causal intervention to study LLMs as organized cognitive systems [9] [13].This bridges brain-inspired artificial intelligence and mechanistic interpretability beyond architectural analogy.
- Limitations and future directions: Evaluations on representative Gemma and LLaMA models spanning pretrained and instruction-tuned stages indicate that the recovered organization is not specific to one checkpoint or training regime [24] [25].Important boundaries remain: multimodal systems differ in how sensory streams are encoded, aligned and integrated, motivating extensions to cross-modal, dynamic and causal organization.
Methods
NeuroCogMap organizes LLM internal activity into functional parcels, cognitive descriptions, capability mappings, and a hierarchy of cognitive operations. The methods combine literature-derived capability curation, SAE-based activation extraction and clustering, atlas selection, and predefined analysis units across model, pathology, cortical, and participant-level evaluations.
- Framework construction: The framework was built by curating capabilities and evaluation datasets, extracting SAE activations, clustering features into functional parcels, and linking parcels to descriptions, capabilities, and cognitive operations.These four linked representations define the NeuroCogMap atlas.
- Models and activation extraction: Analyses used Gemma2-2B, Gemma2-9B-IT, and Llama-3.1-8B, with inputs represented as question–answer pairs and activations extracted over answer spans.The methods distinguish pretrained base models from the instruction-tuned model and use sentence-level SAE activations as the primary atlas observations.
- Feature representation: SAE feature space was used instead of raw neuron space because polysemantic neurons are difficult to assign to stable functional roles, while sparse latent directions support parcellation, cross-model comparison, and intervention.A neuron-based baseline performed worse than the SAE-derived parcel representation on held-out function-predictability datasets.
- Atlas construction: The atlas retained selective, task-responsive features, normalized and reduced their profiles, then clustered them using functional similarity with layer-distribution regularization.Selectivity thresholds retained quantiles of 0.5 for Gemma2-2B and 0.8 for Gemma2-9B-IT and Llama-3.1-8B; the layer regularization weight was λlayer = 0.01.
- Atlas selection: For the primary Gemma2-2B atlas, an integrated clustering-and-interpretability score selected 270 parcels from candidate values of 10–300.The score combined gap-statistic clustering quality, LLM-assisted functional-description quality, and non-redundancy estimates; stability was tested across 16 hyperparameter settings.
- Evaluation design: The evaluation design defined analysis units by task: generated responses for pathology, parcels or capabilities for atlas validation, cortical parcels for human encoding, and participants for participant-level analyses.Decision-model discovery used held-out Psych-101 participants under the Centaur open-loop protocol, with the two-step task serving as the discovery paradigm.
Supplementary Information · Aggregate category analysis of pathology-related activation differences · Social bias reflects misrouting of socially salient representations
NeuroCogMap shows that social bias reflects misrouting of socially salient representations into factual retrieval, whereas fair responses recruit comparison, uncertainty-sensitive reasoning and normative control. Across pathologies, hallucination and bias preferentially perturb belief-related rather than behavioural-control-related systems.
- Aggregate category analysis of pathology-related activation differences: Hallucination and bias produced larger absolute activation differences in belief-related parcels and capabilities than in behavioural-control-related systems.The analysis compared pathological with non-pathological examples across the two aggregate categories.
- Social bias reflects misrouting of socially salient representations: Biased responses showed abnormal coupling between demographic representations and knowledge-retrieval pathways, unlike fair responses that preferentially recruited structured comparison and neutral inference.The disability-domain pattern indicates that clinically or socially associated category information was treated as relevant evidence rather than as a stereotype-laden cue requiring control.
- Social bias reflects misrouting of socially salient representations: Bias-enhanced connectivity linked Celebrity Knowledge Retrieval, Wikipedia Fact Retrieval and Fact Extraction Module to Demographic Comparison, with Δ𝑤 values of −8.68, −8.68 and −6.91, respectively.The corresponding P values were 1.8 × 10−24, 9.3 × 10−35 and 2.6 × 10−17.
- Social bias reflects misrouting of socially salient representations: Fair responses recruited stronger comparison, enumeration and uncertainty-sensitive routes, including Number-Retrieval Module and Locative Retrieval Module control of Uncertainty Reasoner.Fair-response connectivity also involved stronger Factual Retrieval Module–Neutral Inference Module and Neutral Inference Module–Demographic Comparison coupling in the reported contrasts.
- Social bias reflects misrouting of socially salient representations: Fair responses preferentially activated evaluative and alignment-related capabilities, including Causal reasoning, Induction and inference, Analytical thinking, Ethical reasoning and Alignment.Reported activation differences were Δ𝑐=19.92, 13.63, 11.28, 10.27, 8.22 and 8.18, respectively.
- Social bias reflects misrouting of socially salient representations: At the hierarchy level, bias increased Perceptual Access and Attentional Gating, whereas fair responses shifted toward Abstract Reasoning, Meta-Cognitive Control, Situated Application and Social Interaction.This indicates an imbalance between early access to socially salient categories and higher-order processes evaluating, comparing and regulating their relevance.
- Social bias reflects misrouting of socially salient representations: Social bias is a systems-level representational pathology: socially salient category information becomes over-coupled to factual retrieval while comparison, uncertainty-sensitive reasoning and normative evaluation are insufficiently recruited.This distinguishes bias from hallucination, although both involve representational failure.
Mechanism-guided detection and mitigation of social bias · Sycophancy reflects distributed weakening of independent judgment
NeuroCogMap detects social bias near ceiling and enables dataset-specific fairness gains through mechanism-guided intervention. Sycophancy instead reflects distributed weakening of independent evaluative control, disrupting contradiction monitoring, normative reasoning and self-correction in favour of socially aligned responses.
- Mechanism-guided detection and mitigation of social bias: 0.992 mean AUROC on Gemma2-9B-IT exceeded hidden probing (0.991), perplexity (0.733) and attention probing (0.569) for detecting biased outputs.These results support structured parcel-level pathology over generic salience or surprisal measures.
- Mechanism-guided detection and mitigation of social bias: NeuroCogMap-guided intervention increased fairness accuracy on Gemma2-2B from 58.3% to 90.1% on BBQ-Age and from 58.2% to 93.7% on BBQ-Gender.It also increased BBQ-Nationality performance from 60.4% to 65.6% and BBQ-Disability performance from 56.2% to 88.8%.
- Mechanism-guided detection and mitigation of social bias: Bias altered socially salient information routing into retrieval and normative-control pathways, whereas hallucination disrupted factual retrieval, evaluation and cross-system coordination.These multilevel signatures were predictive and steerable under parcel intervention.
- Sycophancy reflects distributed weakening of independent judgment: Independent responses recruited stronger control-oriented coupling among Legal-Moral Evaluation, Decision-Initiation Module and Moral Lesson Detector, including Δw= 0.28, P= 4.2 × 10−15.They also showed stronger Causal Explanation Analysis coupling and inhibitory control linking normative evaluation and decision initiation to Self-Refusal Detection.
- Sycophancy reflects distributed weakening of independent judgment: Sycophantic responses strengthened obedience and self-referential social pathways while disrupting contradiction-sensitive retrieval, indicating a shift in control state rather than merely stronger agreement.The pattern included weaker Self-Pride Recognition–Harassment Detection Module and Narrative Media Comprehension–Listening and Obedience Processing links, each with Δw= −0.26 and −0.20, respectively.
- Sycophancy reflects distributed weakening of independent judgment: Independent responses more strongly activated Social Norm Reasoning (Δa= 13.98, P= 4.2 × 10−32), Combinatorial Reasoning (Δa= 12.50, P= 0.0053) and Legal-Moral Evaluation (Δa= 4.51, P= 6.4×10−23).These parcel-level differences indicate that independence is actively maintained through social-norm evaluation and related reasoning processes.
- Sycophancy reflects distributed weakening of independent judgment: Independent responses showed stronger Fairness, Ethical reasoning, Judgment, Safety and Verification, whereas sycophantic responses were relatively stronger in Common sense reasoning and Logical reasoning.This capability dissociation distinguishes evaluative independence from socially aligned response selection.
- Sycophancy reflects distributed weakening of independent judgment: Sycophancy attenuated independent processing across all four NeuroCogMap layers, with the largest separation in Perceptual Access and Attentional Gating and Situated Application and Social Interaction.Unlike refusal failure’s procedural execution release, sycophancy reflects distributed weakening of conflict monitoring, normative evaluation and self-corrective judgment.
Mechanism-guided detection and mitigation of sycophancy … Hallucination detection baselines
NeuroCogMap-derived pathology signatures supported sycophancy detection and intervention, while the merged analyses defined pathology datasets and standardized comparisons against uncertainty, consistency, and generic internal-state baselines. Hallucination baselines ranged from token-level uncertainty to semantic consistency and hidden-state probing, each with distinct interpretive limitations.
- Mechanism-guided detection and mitigation of sycophancy: NeuroCogMap detected sycophancy above Perplexity and Self-Consistency on Gemma2-2B, reaching AUROC 0.685 for Answer and 0.646 for Feedback.Perplexity achieved 0.571 and 0.543, while Self-Consistency achieved 0.651 and 0.539, respectively.
- Mechanism-guided detection and mitigation of sycophancy: NeuroCogMap-guided steering reduced sycophancy on Gemma2-2B from 28.4% to 26.4% in Answer and from 59.1% to 57.3% in Feedback.Lower sycophancy rates indicated better mitigation.
- Mechanism-guided detection and mitigation of sycophancy: Refusal failure and sycophancy were characterized as behavioural-control pathologies: failures of action gating, conflict monitoring, normative evaluation, and self-correction rather than lost general competence.Hallucination and social bias were instead treated as representational pathologies involving distorted factual, evidential, or demographic representations.
- Pathology datasets and representative task patterns: The pathology datasets paired TruthfulQA and NQ-Open for hallucination, BBQ for social bias, AdvBench and JBB-Behaviors for refusal failure, and Sycophancy-Eval Answer and Feedback for sycophancy.These tasks respectively tested misconception suppression or retrieval, demographic inference, harmful-instruction refusal, factual judgment under pressure, and user-flattering evaluative shifts.
- Shared evaluation protocol: All pathology detectors used binary normative-versus-pathological classification, separating supported from hallucinated responses, successful refusal from compliance, independent from sycophantic responses, and fair from biased responses.Scalar baselines were evaluated with held-out AUROC, while feature-vector baselines used lightweight within-fold classifiers and held-out probabilities.
- Hallucination detection baselines: Hallucination comparisons covered token uncertainty, sequence likelihood, semantic uncertainty, self-consistency, and generic hidden representations without NeuroCogMap’s parcel, capability, and circuit organization.The baseline families included Entropy, LN Entropy, Perplexity, Semantic Entropy, SelfCheckGPT, and Hidden Probing.
- Hallucination detection baselines: Hallucination baselines had complementary limitations: uncertainty and likelihood measures can miss high-confidence or fluent hallucinations, semantic and self-consistency methods require multiple generations, and hidden probing lacks functional or circuit-level attribution.Semantic Entropy groups sampled responses by meaning, SelfCheckGPT measures cross-sample inconsistency, and Hidden Probing decodes labels from generic activations.
Refusal-failure detection baselines · Sycophancy detection baselines
The baselines detect refusal failure through early output likelihoods, perturbation sensitivity, or perplexity, and detect sycophancy through response consistency, user-cue attention, or likelihood comparisons. These methods target prompt sensitivity, local user-belief weighting, or output separability rather than NeuroCogMap’s structured internal-control organization.
- Refusal-failure detection baselines: Refusal-failure baselines classify successful refusals versus refusal failures or jailbreak successes using prompt likelihood, early output distributions, or perturbation sensitivity.The baselines were designed to test whether harmful compliance can be detected without modeling internal control organization.
- Refusal-failure detection baselines: Logits SVM aggregates early next-token probabilities from one forward pass and uses an RBF-SVM to classify refusal failure.Its efficiency comes from avoiding repeated generations, but it treats failure as separability in early output distributions rather than a structured control breakdown.
- Refusal-failure detection baselines: SmoothLLM generates randomized prompt perturbations and detects adversarial prompts through response-cluster stability or entropy.Benign prompts are expected to remain stable, whereas adversarial prompts are expected to produce less stable outputs.
- Refusal-failure detection baselines: Perplexity flags high-likelihood-unusual prompts as potentially pathological, but it can miss fluent harmful instructions and overflag benign rare prompts.This baseline relies on distributional unusualness rather than internal safety or refusal mechanisms.
- Sycophancy detection baselines: Sycophancy baselines compare independent responses with responses shifted toward the user’s stated belief, preference, or self-presentation [96] [33].Because sycophancy detection has less standardization, the baselines were assembled from existing analyses and consistency-based methods.
- Sycophancy detection baselines: Self-Consistency samples repeated answers under the same user-conditioned prompt and detects sycophancy through low semantic consistency or high cluster entropy [33] [96].The method adapts hallucination-detection consistency measures to user-guided response shifts rather than factual unsupportedness alone.
- Sycophancy detection baselines: User-Conditioned Attention measures attention to user-belief or opinion cues, using cue-attention features and a lightweight classifier to distinguish independent from sycophantic responses.It tests local over-weighting of the user’s stance but does not model broader circuit-level or capability-level control changes.
- Sycophancy detection baselines: Sycophancy perplexity features compare likelihoods for correct, user-suggested incorrect, generated, and neutral responses under user-conditioned or feedback settings.The feedback setting uses an unguided neutral response as the reference, including absolute and signed likelihood differences.
Social-bias detection baselines
Social-bias detection baselines tested whether biased responses could be identified from demographic-attention patterns, generic hidden-state representations, or likelihood asymmetries among stereotype-relevant alternatives. They contrasted normative fair or uncertainty-preserving responses with pathological biased responses, while probing distinct but limited signals of bias.
- Social-bias detection baselines: The baselines compared unbiased, fair or uncertainty-preserving responses against biased responses using attention, hidden-state, and perplexity signals motivated by prior bias and fairness work.Together, these methods tested demographic-attention abnormalities, hidden-state label separability, and likelihood asymmetries among stereotype-relevant answer options.
- Attention Probing: Attention Probing used answer-to-attribute attention, including target, counter-attribute, and difference features, followed by logistic regression to predict biased responses.The method tested whether social bias was locally reflected in attention allocation to demographic or protected-attribute spans.
- Hidden Probing: Hidden Probing aggregated middle- and late-layer hidden states into response-level representations and used them to estimate the probability of a biased response.It tested whether bias labels were decodable from generic internal states, without identifying responsible functional parcels or capabilities.
- Perplexity: Perplexity compared likelihoods for substantive stereotype-relevant alternatives after removing uncertainty or abstention options only for feature construction, while retaining original labels for AUROC evaluation.This baseline tested whether biased responses were accompanied by likelihood asymmetries, without identifying how demographic information influenced inference.
Comparison with NeuroCogMap … Capability-associated parcels show hierarchical semantic and structural organization
NeuroCogMap detects multiple safety pathologies more effectively than generic uncertainty, likelihood, logit, and perturbation baselines in Llama-3.1-8B, while parcel construction remains stable across clustering settings. Across models, capability-associated parcels show hierarchical semantic diversification and structural support, linking higher-level integration to a broadly connected lower-level substrate.
- Comparison with NeuroCogMap: NeuroCogMap represents pathology through coordinated deviations in parcel activation, capability recruitment, and circuit connectivity rather than generic output uncertainty or instability.The comparison tests whether pathology is better detected as disruption of internal cognitive organization than as a generic distributional or response-level change.
- Cross-model safety-pathology detection and intervention in Llama-3.1-8B: NeuroCogMap also ranked highest for hallucination detection on NQ-Open and TruthfulQA, reaching AUROCs of 0.691 and 0.663, respectively.The signature generalized across open-domain factual answering and truthfulness-oriented evaluation, with the largest margin on TruthfulQA.
- Cross-model safety-pathology detection and intervention in Llama-3.1-8B: 0.984 and 0.953 AUROC: NeuroCogMap most clearly detected refusal failure on AdvBench and JBB-Behaviors, exceeding Perplexity, Logits SVM, and SmoothLLM.These results indicate stronger cross-model detection from structured cognitive signatures than from prompt likelihood, early-output logits, or perturbation sensitivity.
- Cross-model safety-pathology detection and intervention in Llama-3.1-8B: Sycophancy detection was modest but generally exceeded Perplexity and Self-Consistency, while social-bias detection reached AUROC 1.00 across four BBQ subsets under near-ceiling separability.NeuroCogMap was comparable to User-Conditioned Attention for sycophancy, and the BBQ values were reported as supplementary rather than a main visual comparison.
- Cross-model safety-pathology detection and intervention in Llama-3.1-8B: Parcel intervention improved hallucination accuracy across MedHallu, NQ-Open, and TruthfulQA, while social-bias effects were heterogeneous across BBQ subsets.Hallucination accuracy rose from 78.6% to 84.7%, 42.9% to 46.1%, and 26.9% to 29.9%; social-bias accuracy decreased on BBQ-Age, changed minimally on BBQ-Nationality, and improved on BBQ-Gender.
- Parcel construction is stable across clustering hyperparameters: 0.786 mean similarity across 16 parcellation settings showed stable parcel construction, with within-threshold iteration changes nearly identical and all solution pairs exceeding permutation-null similarity.The overall range was 0.642–1.000, within-threshold mean similarity was 0.999, and all one-sided permutation tests gave P≤0.001.
- Cross-model NeuroCogMap atlas and parcel–capability mapping summaries: Cross-model atlas summaries compare NeuroCogMap’s functional and cognitive atlas properties and its parcel–capability mapping properties across models.Table 2 summarizes functional and cognitive atlas properties, while Table 3 summarizes parcel–capability mapping properties using panel names aligned with Fig. 2.
- Capability-associated parcels show hierarchical semantic and structural organization: Higher hierarchy levels showed broader parcel semantics, whereas lower levels occupied more connected structural positions, supporting a division between integrative higher-order cognition and reusable lower-level support.The analysis used semantic dispersion and degree in a thresholded model-specific structural connectome, replicated in Gemma2-2B and Llama-3.1-8B.
Synthetic examples for hierarchy-dependency validation · Psych-101 cognitive-task datasets used in Fig. 6b
The paper validates hierarchy dependence with synthetic in-context, parametric, and two-hop reasoning examples, then uses five Psych-101 datasets to predict participant-level behavioural fit across cognitive domains. The datasets cover memory, categorization, intertemporal choice, and sequential learning tasks drawn from established experiments.
- Synthetic examples for hierarchy-dependency validation: Synthetic benchmarks tested whether lower-level access and representation processes support higher-level reasoning across in-context, parametric, and mixed two-hop conditions.Profiles contained names and five attributes, enabling controlled questions that required retrieving facts or relating them across hops.
- Synthetic examples for hierarchy-dependency validation: Two-hop parametric examples required deriving a relational answer from biography-style facts, with reasoning traces explicitly connecting intermediate facts to the final answer.For example, identifying who shared a person’s birth date required first retrieving that date and then comparing people associated with it.
- Synthetic examples for hierarchy-dependency validation: Mixed reasoning conditions combined contextual and profile-derived parametric facts to test whether complementary lower-level supports improved higher-level relational reasoning.Final evaluation used temperature 0, top-p = 1, maximum length 50, and exact gold-answer-string containment for correctness.
- Psych-101 cognitive-task datasets used in Fig. 6b: Fig. 6b analysed five Psych-101 participant-level datasets to test whether NeuroCogMap activations predict variation in behavioural fit between a language model and individual humans.The domains were episodic long-term memory, multi-attribute decision-making, Shepard categorization, intertemporal choice, and drifting four-armed bandit learning.
- Psych-101 cognitive-task datasets used in Fig. 6b: The episodic memory task involved repeated study, arithmetic distraction, and recall cycles, with coloured borders distinguishing intentional from incidental learning.Participants also made semantic judgements during study.
- Psych-101 cognitive-task datasets used in Fig. 6b: The Shepard categorization task required learning feedback-based rules for geometric objects varying in features such as shape, size, and colour.Participants judged each object and received feedback that supported learning across the sequence.
- Psych-101 cognitive-task datasets used in Fig. 6b: The intertemporal-choice task contrasted smaller immediate with larger delayed monetary outcomes, including both gains and losses to probe discounting and framing.The drifting four-armed bandit task required sequential reward updating while balancing exploitation and exploration as option values changed over time.
Human evaluation of LLM-assisted NeuroCogMap components · NeuroCogMap supplementary atlas and dataset resources
Human audits supported the reliability of NeuroCogMap’s LLM-assisted parcel annotations, pathology labels and cross-model correspondences, while supplementary tables document its hierarchy, taxonomy, datasets, parcel atlas and cross-model mappings.
- Human evaluation of LLM-assisted NeuroCogMap components: Human audits supported NeuroCogMap’s LLM-assisted components, with positive validation of parcel descriptions, pathology labels and cross-model parcel correspondences.The audit covered 60 parcels, 570 pathology responses and 100 matched parcel pairs.
- Human evaluation of LLM-assisted NeuroCogMap components: 93.3% of sampled parcel descriptions were judged pass-or-partial, with high ratings for specificity, faithfulness and coverage and low overreach.Across 60 parcels, specificity averaged 4.24, faithfulness 4.46, coverage 4.41 and overreach 1.71 on 1–5 scales.
- Human evaluation of LLM-assisted NeuroCogMap components: Human and automated pathology labels matched exactly for 417/570 responses (73.2%), while coarse normative-versus-pathological agreement reached 426/515 val…Inter-annotator agreement was substantial (Fleiss’ κ = 0.783), and validation varied by pathology type, with refusal failure showing the strongest agreement at 91.3% exact and 93.8% coarse agreement.
- Human evaluation of LLM-assisted NeuroCogMap components: 77.0% of 100 matched parcel pairs were judged identical or partially similar, supporting most reported Gemma2-2B–Gemma2-9B-IT correspondences.Mean human similarity was 2.30/3, with high reliability (ICC = 0.838); 71/100 pairs also had mean similarity of at least 2.
- Human evaluation of LLM-assisted NeuroCogMap components: Human similarity ratings closely tracked the pipeline’s matching scores, with the combined score correlating at ρ = 0.856 and LLM-classified identical pairs scoring 2.96 versus 1.97 for partially similar pairs.The combined-score correlation had P = 8.0 × 10^-30, and the identical-versus-partial distinction achieved AUROC = 0.850.
- NeuroCogMap supplementary atlas and dataset resources: Supplementary resources document the cognitive hierarchy, capability taxonomy, benchmark resources, datasets used for construction and model-derived atlas summaries.The resources are presented as compact manuscript tables rather than raw data inventories.
- NeuroCogMap supplementary atlas and dataset resources: The supplementary atlas lists interpretable Gemma2-2B parcels spanning retrieval, reasoning, safety, bias detection, language processing and other specialized functions.Examples include comparative reasoning, mathematical reasoning, harassment detection, social-bias detection, temporal reasoning and procedural guidance modules.
- NeuroCogMap supplementary atlas and dataset resources: Additional tables map capabilities to parcels across language models, extending the atlas resources beyond the single-model parcel listings.The cross-model capability-to-parcel mapping is provided in Table 11.
Supplementary prompt templates
This section summarizes the prompt templates and structured input-output formats used for NeuroCogMap annotation, evaluation, and model-discovery analyses. It omits internal file paths and execution-specific metadata while listing the return fields enforced during implementation.
- Supplementary prompt templates: The templates define structured input-output formats for NeuroCogMap annotation, evaluation, and model-discovery analyses.
- Supplementary prompt templates: Internal file paths and execution-specific metadata are omitted from the listed templates.
- Supplementary prompt templates: The listed return fields correspond to the structured-output fields enforced during implementation.
Cognitive atlas parcel annotation prompts · Parcel-function validation and cross-model comparison prompts
NeuroCogMap parcels are annotated from activation samples, keywords, and dataset distributions using neuroscience-inspired functional descriptions. Validation and comparison prompts score interpretability, detect redundancy, predict activation, assess intervention alignment, and compare parcels across models.
- Cognitive atlas parcel annotation prompts: Parcel annotations use activation samples, ranked keywords, and dataset distributions to infer each parcel’s processed information, concise function name, detailed description, and role in the model.High-activation examples include the question, answer, activated sentence, activation strength, and dataset.
- Parcel-function validation and cross-model comparison prompts: Functional descriptions receive 0–1 scores for specificity, example and keyword coverage, mechanistic clarity, neuroscience-style interpretability, and avoidance of vague or circular explanations.The evaluator returns both a score and a brief rationale.
- Parcel-function validation and cross-model comparison prompts: Redundancy checks compare parcel descriptions by their core information-processing roles, ignoring superficial wording differences and returning a binary redundancy judgment, similarity score, and rationale.The comparison uses function names, descriptions, and keywords for both parcels.
- Parcel-function validation and cross-model comparison prompts: Activation-ranking prompts estimate each candidate parcel’s 0–1 activation probability for an input example using its function name, description, and keywords, then return a ranked list with rationales.Candidates are supplied with their parcel descriptions, and the output includes parcel IDs and activation probabilities.
- Parcel-function validation and cross-model comparison prompts: Intervention-alignment evaluation scores original and post-intervention responses from 0–1 and classifies alignment change as increased, decreased, or no_change.The judgment evaluates whether responses align with the specified parcel function and includes a rationale.
- Parcel-function validation and cross-model comparison prompts: Cross-model parcel comparisons assess functional positioning, concept overlap, dataset-distribution similarity, and semantic relatedness to produce a similarity score and an Identical, Partially Similar, or Different label.Each parcel is represented by its model, function name, description, keywords, and dominant datasets.
Parcel-to-capability semantic templates
NeuroCogMap uses structured semantic templates to represent capabilities and parcels, then interprets their normalized similarity as evidence of functional alignment. High similarity suggests semantic consistency and potential support, whereas low or negative similarity indicates weak or opposing alignment.
- Template structure: Capability templates encode definitions, cognitive alignment, and manifestations in LLMs, while parcel templates encode functions, descriptions, and roles in large models.These paired fields provide the semantic information used to relate capabilities to parcels.
- Similarity interpretation: Capability–parcel similarity measures how closely a parcel’s function matches a capability’s semantic definition.The similarity score is interpreted after normalization.
- Similarity interpretation: High similarity indicates semantic consistency and possible support for the capability, whereas low or negative similarity indicates weak or opposing alignment.The score is therefore an interpretive signal rather than a direct behavioral measurement.
Pathology evaluation and classification prompts
The section specifies structured prompts for evaluating hallucination, bias, jailbreak/refusal failure, and sycophancy, alongside a classifier that labels NeuroCogMap units by belief and control functions.
- Hallucination evaluation: Hallucination evaluation checks whether a model answer matches a reference in meaning and factual content, allowing paraphrases but rejecting unsupported factual additions.The prompt returns correctness, the matched reference, confidence, and a rationale.
- Bias evaluation: Bias evaluation first identifies uncertainty, then maps answers to BBQ options and labels them no_bias, bias, or antibias according to stereotype alignment.The output includes the uncertainty response, mapped option, label, confidence, and rationale.
- Jailbreak and refusal-failure evaluation: Jailbreak evaluation assesses refusal, harmful content, and actionable detail, then reports whether the harmful request produced a successful jailbreak.The prompt separately returns refusal, harmful-content, actionable-detail, jailbreak-success, confidence, and rationale fields.
- Sycophancy evaluation: Sycophancy evaluation tests whether a preference-conditioned response shifts toward user flattery relative to a neutral response, retaining only judgments consistent across presentation orders.It returns sycophantic shift, directional consistency, confidence, and rationale.
- Belief/control classification: The belief/control classifier assigns NeuroCogMap units to belief-related, control-related, mixed, neutral, or unknown categories based on their described functions.Belief-related functions include semantic representation and knowledge integration, whereas control-related functions include inhibition, arbitration, safety regulation, and goal control.
Human cortical description and human–LLM correspondence prompts · Model-discovery agent prompts
The prompts operationalize human cortical parcel description, human–LLM functional correspondence scoring, and NeuroCogMap-guided discovery of cognitive mechanisms and executable model extensions. They combine behavioural evidence with NeuroCogMap activations and preserve comparable task interfaces and search budgets across discovery branches.
- Human cortical description and human–LLM correspondence prompts: Human cortical parcels are described from Cognitive Atlas terms with signed z-scores, producing a concise function name, 100–200-word description, and brief brain-function role.The prompt takes a parcel name and associated term profile as inputs.
- Human cortical description and human–LLM correspondence prompts: Human–LLM correspondence is judged by comparing a human parcel description with candidate NeuroCogMap parcel descriptions across major cognitive domains.The output includes a 0–1 score, match label, best parcel, and rationale.
- Model-discovery agent prompts: The reasoning agent receives behavioural traces, AIC comparisons with the Dual-systems Model, residual summaries, and labels identifying regimes where the LLM fits better.These inputs support mechanism discovery in selected two-step tasks.
- Model-discovery agent prompts: NeuroCogMap-guided discovery additionally uses parcel and capability activations, cognitive descriptions, and contrasts with regimes where the baseline model performs as well or better.These inputs provide internal representational evidence beyond behaviour-only discovery.
- Model-discovery agent prompts: The reasoning agent proposes cognitive mechanisms explaining why the LLM simulator better fits human behaviour than the baseline cognitive model.Each proposal links behavioural evidence, available NeuroCogMap evidence, a formal model component, and a rationale.
- Model-discovery agent prompts: The model-writing agent converts proposed mechanisms into executable extensions while preserving the baseline task interface and participant-level choice-data fitting.The branch uses the same search budget as behaviour-only discovery.
- Model-discovery agent prompts: Model outputs include candidate code, parameters, mechanism-to-component mappings, and compatibility notes for the baseline fitting pipeline.This output format makes proposed extensions implementable and comparable within the existing workflow.