Source-linked AI summary
NeuroQuery: comprehensive meta-analysis of human brain mapping
Jérôme Dockès, Russell Poldrack, Romain Primet, Hande Gözükan, Tal Yarkoni, Fabian Suchanek, Bertrand Thirion, Gaël Varoquaux
TL;DR
Brain-imaging evidence spans many concepts and inconsistent terms, while conventional meta-analysis mainly tests publications associated with frequent single terms. NeuroQuery instead uses a multivariate predictive model to map arbitrary text to likely brain locations, learning from 7 547 terms across 13 459 publications. The resulting tool supports hypothesis generation and analysis priors, but individual predictions are not definitive statistical conclusions.
Problem
Brain-imaging meta-analysis is limited by heterogeneous concepts and terminology, while existing statistical approaches mainly handle frequently occurring single terms.
Method
NeuroQuery uses semantic associations and supervised machine learning to predict brain maps from arbitrary text queries using a corpus of full-text neuroimaging publications.
Results
NeuroQuery captures 7 547 neuroscience terms across 13 459 neuroimaging publications and predicts brain locations for single terms, combinations, and detailed descriptions.
Takeaways & Limitations
NeuroQuery extends meta-analysis to less frequently studied concepts and can support hypothesis generation, regions of interest, and formal analysis priors.
Takeaways & Limitations
A specific NeuroQuery prediction may be wrong and does not by itself support definite conclusions because it lacks a statistical test.
Abstract
from arXiv · showhide
Reaching a global view of brain organization requires assembling evidence on widely different mental processes and mechanisms. The variety of human neuroscience concepts and terminology poses a fundamental challenge to relating brain imaging results across the scientific literature. Existing meta-analysis methods perform statistical tests on sets of publications associated with a particular concept. Thus, large-scale meta-analyses only tackle single terms that occur frequently. We propose a new paradigm, focusing on prediction rather than inference. Our multivariate model predicts the spatial distribution of neurological observations, given text describing an experiment, cognitive process, or disease. This approach handles text of arbitrary length and terms that are too rare for standard meta-analysis. We capture the relationships and neural correlates of 7 547 neuroscience terms across 13 459 neuroimaging publications. The resulting meta-analytic tool, neuroquery.org, can ground hypothesis generation and data-analysis priors on a comprehensive view of published findings on the brain.
1 Introduction: pushing the envelope of meta-analyses
Brain-imaging meta-analysis must integrate a large, heterogeneous literature despite inconsistent terminology and concepts. NeuroQuery addresses these constraints by predicting brain maps from contextual text rather than testing publications selected by exact terms.
- Motivation: More than 6 000 neuroimaging publications appear annually, making manual aggregation of findings across diverse behaviors and protocols impractical.Individual studies often lack sufficient statistical power to establish fully trustworthy results.
- Terminology and coverage: Automated meta-analysis remains constrained by inconsistent terminology: only 30% of terms in one neuroscience ontology or tool are shared by another.Related terms can also produce markedly different meta-analyses.
- Terminology and coverage: Standard methods cannot automatically identify equivalent functional contrasts or interpolate evidence for rare diseases and combinations of mental processes.Consequently, they cannot answer questions that cannot be formulated simply.
- NeuroQuery's approach: NeuroQuery reframes meta-analysis as out-of-sample prediction, mapping descriptions of experiments, diseases, or psychological concepts to likely anatomical structures.This complements hypothesis-testing methods such as ALE, MKDA, and NeuroSynth.
- NeuroQuery's approach: NeuroQuery combines multiple terms contextually and learns from 13 459 full-text publications to predict brain locations for queries ranging from single words to detailed descriptions.Its vocabulary contains 7 547 neuroscience terms, and the model uses semantic relations between terms.
2 Results: The NeuroQuery tool and what it can do.
NeuroQuery combines semantic relationships among neuroscience terms with supervised text-to-brain regression to generate maps for arbitrary queries. It generalizes to unseen term combinations and difficult concepts, predicts coordinates in new studies, and agrees with several established meta-analytic references, while its maps do not support simple null-hypothesis thresholding.
- Model and pipeline: Semantic smoothing expands a query with related terms before projection onto reliably encoded keywords and mapping into brain space.The procedure helps address rare, polysemic, or weakly brain-associated expressions while retaining higher weight for terms explicitly present in the query.
- Scope and interpretation: Unlike standard meta-analysis maps, NeuroQuery maps cannot be thresholded to reject a simple null hypothesis because they have different meanings and uses.The model is intended as a predictive mapping approach rather than an in-sample statistical summary.
- Generalization to new queries: NeuroQuery predicts useful maps for unseen combinations by additively combining maps for terms whose joint occurrence was excluded from training.For the distance–color example, predicted maps included structures associated with distance perception and an additional visual region associated with color perception; performance decreased only slightly for unseen versus previously seen pairs across 1 000 random pairs.
- Quantitative evaluation: More than 72% of the time, NeuroQuery’s predicted map was closer to the correct held-out study map than to a randomly selected negative example.This evaluation used 16-fold cross-validation on unseen neuroimaging studies and compared predicted maps with reported coordinates.
- Agreement with reference maps: NeuroQuery predictions aligned with curated IBMA maps at median AUC 0.80, anatomical atlas regions at median AUC 0.98, and NeuroSynth maps at median AUC 0.90.The IBMA comparison reached median AUC 0.88 for NeuroQuery after manually reformulating labels to NeuroSynth vocabulary; the NeuroSynth comparison used frequent-enough terms.
3 Discussion and conclusion
NeuroQuery reframes automated meta-analysis as prediction, mapping arbitrary neuroscience queries to brain maps through learned semantic and neural associations. The tool broadens coverage beyond frequent single terms while requiring users to treat predictions as probabilistic aids rather than definitive statistical conclusions.
- Approach: NeuroQuery predicts brain maps for arbitrary text queries by combining semantically related keyword maps through supervised learning.Its model uses semantic associations to select reliably mapped keywords, then transforms their weighted representation into a brain map.
- Related work: Unlike topic-modeling approaches, NeuroQuery maps individual terms and uses term co-occurrences to infer semantic links rather than reducing concepts to latent topics.It uses supervised encoding rather than decoding from brain locations to neuroscience terms.
- Limitations: Specific NeuroQuery predictions can be wrong despite good average performance and do not support definite conclusions because they lack a statistical test.The authors recommend using the tool to generate hypotheses and data-analysis priors rather than as a replacement for inferential meta-analysis.
- Limitations: Semantic smoothing can produce misleading maps when a query is represented indirectly by distant or overrepresented associations.The paper gives “agnosia” and “ADHD” as examples and recommends inspecting associated keywords, individual maps, weights, and the final combination.
- Broader considerations: Meta-analysis cannot by itself correct biases in the primary literature or establish that a brain structure is selective for a mental condition.Selectivity claims require an explicit statistical model of reverse inference, while NeuroQuery accounts for varying baseline reporting across brain locations.
- Practical scope: NeuroQuery is especially useful for rare terms, multi-term concepts, and fully automated queries that would otherwise require manual curation.It provides statistical maps for queries ranging from seldom-studied terms to free-text experimental protocols.
A.1 A new dataset
The NeuroQuery dataset expands term-study information beyond abstracts by assembling full-text neuroimaging articles from structured sources. This provides more occurrences and associations for rare terms, while the source selection can introduce journal-availability bias.
- Dataset scope: NeuroSynth contains term frequencies for 3 228 terms based on study abstracts, whereas NeuroQuery uses a corpus of 13 459 full-text studies.The NeuroSynth dataset also contains 448 255 unique locations for 14 371 studies.
- Rare terms: In the NeuroQuery corpus, “huntington disease” occurs in 32 abstracts and “prosopagnosia” in 25, leaving such terms with limited meta-analytic power.Full text exposes many more term occurrences and associations between terms and studies.
- Collection: The authors downloaded around 149 000 full-text journal articles related to neuroimaging from PubMed Central and Elsevier APIs.The sources were chosen because they provide many articles in a structured format.
- Scope boundary: The dataset may have selection bias because some scientific journals, mostly paid journals, are unavailable through the chosen article sources.Articles are converted to JATS XML, validated, and filtered to retain titles, keywords, abstracts, and relevant article-body sections.
A.3 Coordinate extraction
NeuroQuery extracts brain-activation coordinates and article text into standardized representations for large-scale analysis. The pipeline uses coordinate pooling and density estimation while treating all coordinates as MNI-space data.
- Data extraction: Tables are normalized into matrix-shaped XHTML representations before extracting article information.Cells spanning rows or columns are split and their contents duplicated during normalization.
- Coordinate handling: All extracted coordinates are treated as Montreal Neurological Institute coordinates, including coordinates reported in Talairach space.This uniform treatment is adopted despite the two coordinate systems differing by up to about 1 cm.
- Limitations: Improved Talairach-to-MNI handling is identified as a direction for improving the dataset.The authors consider uniform treatment acceptable as a first approximation because coordinate differences are generally comparable to smoothing-kernel size.
- Coordinate handling: Coordinates from all article tables are pooled and converted into brain activation-density maps with Gaussian kernel density estimation.The selected kernel has a Full Width at Half Maximum close to 9mm.
- Coordinate handling: The normalized density of peak coordinates reduces dependence on the number of contrasts and other analytic choices affecting reported-coordinate counts.Reported coordinates can range from fewer than a dozen to several hundred per article.
- Text representation: The text representation uses TFIDF features from manually curated neuroscience vocabularies and ontologies.These vocabularies restrict features to relevant neuroscience terms and meaningful collocations while reducing dimensionality.
A.6 Summary of collected data
The collected NeuroQuery resources combine a large full-text neuroscience corpus, extracted activation coordinates, and a 7 547-term vocabulary. Compared with NeuroSynth, the corpus contains substantially more text and term-study associations, while coordinate extraction is reported as less noisy.
- Collected resources: The dataset contains over 149K full-text neuroscience articles, over 418K peak coordinates from more than 13.5K articles, and 7 547 neuroscience terms.Each vocabulary term occurs in at least 6 articles with extracted coordinates.
- Comparison with NeuroSynth: The NeuroQuery corpus contains 20 times more raw text than NeuroSynth and over 5.5M unique-term occurrences.Restricting to NeuroSynth’s vocabulary still yields over 3M term-study associations, 4.6 times more than NeuroSynth.
- Comparison with NeuroSynth: NeuroQuery’s coordinate extraction reduced articles with incorrect coordinates by a factor of 7 and articles with missing coordinates by a factor of 3 relative to NeuroSynth.The comparison used manually annotated articles where the two extraction systems disagreed.
- Data sharing: The vocabulary, extracted coordinates, and term-occurrence counts are freely available, but the full article texts cannot be shared.
B Methodological details
NeuroQuery represents article text with TFIDF features and predicts voxelwise activation density using adaptive ridge regression. Feature selection and test-time smoothing focus the model on informative terms while reducing computation and supporting interpretable map thresholds.
- Text representation: Each study is represented by 7 547 neuroscience terms or phrases using TFIDF features derived from a fixed curated vocabulary.The representation averages normalized TFIDF vectors from the title, abstract, full text, and keywords.
- Voxelwise regression: The model regresses activation density at each voxel on study-level TFIDF descriptors using ridge regression.The design matrix contains study features, while the dependent variables represent activation density across brain voxels.
- Adaptive regularization: The reweighted feature-selection procedure addresses sparse study features and the need to select terms across approximately 28 000 voxel outputs.It begins with a uniformly regularized ridge fit and uses brain-map information to reweight features.
- Feature selection: Terms with low squared map norms are discarded, while the remaining features are reindexed for the final model.A small margin ε = 0.001 helps avoid division by zero, and the practical feature set contains about 200 terms.
- Prediction and interpretation: Rescaled predictions provide a natural map threshold, with values around ẑ ≈ 3 selecting regions typical of a query.These thresholded regions can support region-of-interest analysis.
- Computational efficiency: Feature selection keeps about 200 features, reduces prediction-time computation, and supports deployment of the online NeuroQuery tool.Its computational cost is two ridge regressions, with the second using the smaller selected feature set.
- Test-time smoothing: The smoothing matrix mixes the identity matrix with term associations using α = 0.1, preserving greater weight for query terms than expansion terms.This design prevents query expansion from degrading predictions when terms are already well encoded.
C.1 Example Meta-analysis results for the RSVP language task from the IBC dataset.
The example meta-analysis reports the studies included for the “Read pseudo-words vs consonant strings” language contrast. The listed GingerALE analysis comprises 29 studies.
- Study set: The GingerALE meta-analysis for “Read pseudo-words vs consonant strings” includes 29 studies.
C.2 NeuroQuery performance on unseen pairs of terms
NeuroQuery is evaluated on studies containing term pairs excluded from training, testing whether it can predict maps for combinations never jointly observed. Predictions remain consistent with held-out meta-analytic evidence and related terminology, while comparisons assess agreement with established references.
- Unseen-pair evaluation: 1000 randomly selected term pairs were evaluated by excluding studies containing both terms from training and testing predictions on those left-out studies.Pairs appeared together in more than 500 but fewer than 1000 studies, and the experiment was repeated 1000 times.
- Unseen-pair evaluation: Unseen-pair prediction showed a slight likelihood decrease and a more marked matching-performance decrease relative to random-study cross-validation.The held-out task is harder because test publications containing the same pair are more similar to one another than randomly selected publications.
- Map quality: The predicted maps achieved a median Pearson correlation of .85 with the average reported-location density in held-out studies.This correlation measures whether predictions resemble the map that a direct meta-analysis of the excluded studies would produce.
- Variable terminology: NeuroQuery produced consistent maps for related mental-arithmetic terms through semantic smoothing.The maps illustrate how information from easier terms such as “calculation” and “arithmetic” can support difficult terms such as “computation” and “addition.”
- Reference comparisons: Against BrainPedia maps, NeuroQuery reached median AUC values of 0.9 with reformulated labels and 0.8 with original labels.NeuroSynth reached 0.8 with reformulated labels and often failed on original labels missing from its vocabulary.
- Reference comparisons: NeuroQuery matched Harvard–Oxford atlas regions and NeuroSynth activation maps with median AUC above 0.9 and 0.90, respectively.The atlas comparison used 18 labels, while the NeuroSynth comparison selected 200 terms with the largest numbers of active voxels.
D Word occurrence frequencies across the corpus
The corpus contains many rare terms, making occurrence frequency central to mapping specialized queries. Full-text articles provide substantially more term–study associations than abstract-only corpora, supporting multivariate prediction for combinations of terms.
- Evaluation context: The reported comparisons include NeuroQuery maps against manually segmented atlas regions and thresholded NeuroSynth activations using Area Under the ROC Curve.AUC is defined as the probability of ranking an active voxel above an inactive one.
- Multivariate queries: Combining terms can use documents containing either term: “face perception” occurs in 413 articles and “dementia” in 1312, while 1703 contain at least one.The multivariate prediction uses documents with a nonzero overlap between the query and document term vectors.
E.1 Details on the Medical Subject Headings
NeuroQuery’s vocabulary is built from neuroscience-relevant branches of MeSH, but rarity and terminology variation constrain which concepts are represented as single phrases. Related multi-term parsing can partially recover some absent variants.
- Vocabulary construction: The vocabulary includes MeSH branches for neuroanatomy, neurological disorders, and psychology relevant to neuroscience and psychology.The included branches are selected from the broader Medical Subject Headings graph.
- Vocabulary limits: Many MeSH terms are excluded because they are too rare or too specific for the corpus.“Diffuse Neurofibrillary Tangles with Calcification” is given as an example of an overly specific term.
- Vocabulary limits: NeuroQuery recognizes “Frontotemporal Dementia” but not several MeSH entry-term variants, including “Dementia, Frontotemporal” and “HDDD1.”These variants may still be represented by combinations of other recognized terms.
- Phrase frequency: Phrase expressions involving several words are typically much rarer than occurrences of at least one constituent word.This frequency difference affects how multi-word concepts are represented in the corpus.
- Atlas terminology: The vocabulary includes labels from 12 atlases, while atlas labels absent from NeuroSynth’s vocabulary are discarded in the anatomical comparison.The supplied material identifies atlas-label inclusion but does not enumerate the 12 atlases.