Source-linked AI summary
BOLD5000: A public fMRI dataset of 5000 images
Nadine Chang, John A. Pyles, Abhinav Gupta, Michael J. Tarr, Elissa M. Aminoff
TL;DR
Human visual-neuroimaging studies have used relatively few images, limiting dataset scale, diversity, and overlap with computer-vision stimuli. BOLD5000 addresses this gap with a slow event-related fMRI study using almost 5,000 images from SUN, COCO, and ImageNet, with analyses validating the dataset. The dataset brings biological and computer vision closer together, while remaining smaller than lifelong visual experience and modern artificial-vision training sets.
Problem
Human neuroimaging studies of visual perception use relatively few images, and their stimuli often lack diversity and overlap with computer-vision datasets.
Method
BOLD5000 uses a slow event-related human fMRI design with almost 5,000 images drawn from SUN, COCO, and ImageNet, analyzed with MRIQC, repeated-stimulus reliability, RSA, and t-SNE.
Results
BOLD5000 provides a large, diverse fMRI dataset with stimulus overlap across standard computer-vision datasets and neural representations showing expected relationships between visual regions and AlexNet layers.
Takeaways & Limitations
BOLD5000 supports closer integration of biological and computer vision by pairing extensive neural measurements with real-world images used in computer vision.
Takeaways & Limitations
Although larger than prior human neuroimaging datasets, BOLD5000 remains small relative to lifelong human visual experience and the million-image datasets used to train modern artificial vision systems.
Abstract
from arXiv · showhide
Vision science, particularly machine vision, has been revolutionized by introducing large-scale image datasets and statistical learning approaches. Yet, human neuroimaging studies of visual perception still rely on small numbers of images (around 100) due to time-constrained experimental procedures. To apply statistical learning approaches that integrate neuroscience, the number of images used in neuroimaging must be significantly increased. We present BOLD5000, a human functional MRI (fMRI) study that includes almost 5,000 distinct images depicting real-world scenes. Beyond dramatically increasing image dataset size relative to prior fMRI studies, BOLD5000 also accounts for image diversity, overlapping with standard computer vision datasets by incorporating images from the Scene UNderstanding (SUN), Common Objects in Context (COCO), and ImageNet datasets. The scale and diversity of these image datasets, combined with a slow event-related fMRI design, enable fine-grained exploration into the neural representation of a wide range of visual features, categories, and semantics. Concurrently, BOLD5000 brings us closer to realizing Marr's dream of a singular vision science - the intertwined study of biological and computer vision.
Background & Summary
BOLD5000 addresses the data gap between computer and biological vision by collecting a large, diverse human fMRI dataset with images overlapping standard computer vision datasets. Its scale and design aim to support closer integration of neural measurements with computational models.
- Human and computer vision share high-level goals, but biological-vision studies typically use far fewer images than modern computer-vision datasets.This mismatch complicates integrating computational models with neuroscience.
- Advanced neuroimaging provides finer-scale measurements of brain activity, yet how visual content maps onto specific brain responses remains an open question.Complex neural activity also makes mechanistic descriptions of high-level visual processes difficult to interpret.
- BOLD5000 collects 5,000 real-world images in a large-scale, slow event-related human fMRI study.The study includes approximately 20 hours of MRI scanning for each of four participants.
- The dataset draws images from SUN, COCO, and ImageNet, creating overlap with commonly used computer-vision datasets.This design connects neural measurements to image domains used in computational vision.
- BOLD5000 uses representational similarity analysis and t-distributed stochastic neighbor embedding visualizations alongside standard fMRI analyses to validate data quality.These analyses are intended to support integration between human and computer vision.
Stimuli
BOLD5000 addresses the limited size, diversity, and cross-domain overlap of visual stimuli in human neuroimaging by presenting almost 5,000 distinct images from standard computer vision datasets. Its stimulus design spans scenes, complex object contexts, and centered objects while controlling image properties for fMRI analysis.
- Cross-domain overlap: Using images from computer vision datasets improves image overlap between neural studies and model training or testing stimuli.The paper identifies this overlap as important for comparing neural and model representations and for using neural data in network training or design.
- Dataset scope: BOLD5000 includes 5,000 real-world images, with 4,916 unique experimental stimuli, greatly expanding the scale of human fMRI visual studies.The study describes this scale as over an order of magnitude larger than nearly all extant human fMRI studies.
- Dataset scope: The stimuli combine SUN-inspired scenes, COCO images containing multiple interacting objects, and ImageNet images focused mainly on single centered objects.These sources cover complementary visual domains, from indoor and outdoor scenes to complex contexts and isolated objects.
- Stimulus preprocessing: The slow event-related design separates responses to individual images, unlike large-scale movie stimuli whose events overlap in time.Overlapping movie events make it difficult to disentangle which stimuli generated particular neural responses.
- Stimulus preprocessing: Image quality controls included square, equal-sized, color images and gray world normalization to make luminance as uniform as possible.The selection process also checked resolution, size, and blurring before normalization.
Experimental Design
The experiment presented thousands of proportionally sampled images across repeated scanning sessions using a slow event-related fMRI design. Participants made valence judgments, while separate localizer runs supported region-of-interest definition.
- Scanning schedule: Four participants viewed 5,254 image trials across 15 functional sessions, with 3,108 trials collected from CSI4.Three participants completed full datasets over 16 sessions; CSI4 completed fewer sessions because of MRI discomfort.
- Stimulus presentation: Each run contained 37 stimuli proportionally matching the overall dataset: roughly one-fifth scenes, two-fifths COCO, and two-fifths ImageNet.Approximately two images per run were repeated, while the remaining 35 were unique.
- Stimulus presentation: A slow event-related design was used to isolate the BOLD signal for each individual image trial.Fixation periods preceded and followed each run, and stimuli were shown sequentially with interstimulus fixation.
- Participant task: For every image, participants performed a valence judgment by indicating whether they liked, felt neutral toward, or disliked it.Responses were made during the post-stimulus fixation interval using an MRI-compatible response glove.
- Functional localizers: Functional localizers used scenes, objects, and scrambled images in block-design runs with a one-back repetition-detection task.Localizer images were non-overlapping with the main experiment’s scene stimuli.
- MRI acquisition: MRI data were acquired on a 3T Siemens Verio scanner using multiband T2*-weighted echo-planar functional imaging plus anatomical and diffusion scans.The acquisition included 69 slices at 2 × 2 mm in-plane resolution and 2 mm slice thickness for functional images.
Data Analyses
The analysis pipeline converted and quality-checked the MRI data, modeled session responses with nuisance regression, and extracted normalized ROI timecourses for subsequent analyses. ROIs were independently localized and sampled using participant-native anatomy.
- Preprocessing: MRI data were converted to BIDS format and assessed with MRIQC before analyses used FMRIPREP-preprocessed data.FMRIPREP performed anatomical correction, skull stripping, surface reconstruction, spatial normalization, and functional preprocessing.
- Analysis pipeline: The pipeline combined motion correction, distortion correction, coregistration, spatial normalization, GLM nuisance regression, and ROI-level analysis.These stages connected standardized preprocessing with participant-specific neural response analyses.
- GLM analysis: Session data entered a general linear model in which nine nuisance variables, including motion and tissue signals, were regressed out.The model was applied to data from sessions containing nine or ten runs.
- ROI definition: Scene-selective ROIs included PPA, RSC, and OPA, defined from scenes-versus-objects-and-scrambled contrasts; LOC was also examined for object selectivity.ROI analyses were conducted individually in native volumetric space using MarsBaR.
- Timecourse extraction: ROI timecourses were averaged across image presentations, with CSI1–CSI3 analyzed at TR3–TR4 and CSI4 at TR3.These windows followed participant-specific peaks in the extracted timecourses.
- Normalization: Extracted ROI data were demeaned voxelwise by subtracting each voxel’s mean across all image presentations from each sample.This normalization was applied after ROI extraction.
Participant Selection
Participants were recruited from Carnegie Mellon’s graduate-student pool for a demanding multisession MRI study. The final sample comprised four healthy adults, three with complete datasets and one with fewer sessions because of discomfort.
- Recruitment and completion: Four Carnegie Mellon graduate students participated, with three completing all 16 MRI sessions and CSI4 completing 10 sessions.CSI4’s reduced participation was attributed to discomfort in the MRI.
- Recruitment and completion: Recruitment favored people familiar with MRI procedures and able to complete the multisession protocol with minimal movement.This selection aimed to reduce effects on data quality during the extended scanning schedule.
- Participant characteristics: Participants were 24–27 years old, right-handed, and reported no psychiatric or neurological disorders or current psychoactive medication use.The sample included one male and three females.
- Ethics: All participants provided written informed consent, and the procedures were approved by Carnegie Mellon University’s Institutional Review Board.The study procedures followed the Declaration of Helsinki.
Code Availability
The complete study scripts and stimulus images are publicly available for download.
- Psychtoolbox Matlab scripts for running the study are available online.
- The complete set of experimental stimulus images is downloadable as BOLD5000_Stimuli.zip.
Data Records
BOLD5000 has been publicly released online, with collected data and analyzed-data stages documented for access.
- BOLD5000 is publicly released online through the project website.The release includes a comprehensive list of collected data and analyzed-data stages.
Technical Validation
Technical validation assesses data quality, stimulus-response isolation, signal reliability, and correspondence between neural representations and computational or image-category structure. The analyses report stable quality measures, reliable repeated-image responses, and interpretable representational organization across visual regions.
- Data Quality: MRIQC evaluates each run with image-quality metrics and reports run-level BOLD averages, standard deviations, and cross-run stability.The reported measures are described as on par with or better than those in comparable studies.
- Repeated Stimulus Images: A one-second stimulus followed by nine seconds of fixation isolates trials, with hemodynamic responses peaking around six seconds and returning to baseline before the next stimulus.
- Repeated Stimulus Images: 4,803 images appeared once, while 113 images were repeated at least three times to assess signal-to-noise reliability.
- Repeated Stimulus Images: Correlation across repetitions of the same image was considerably higher than correlation across different images, indicating reliable trial-level BOLD patterns.
- Representational Similarity Analysis: RSA compares similarity among 4,916 scenes in voxel space with similarity in AlexNet feature space across ROIs and network layers.High-level regions such as PPA correlate more strongly with higher AlexNet layers, whereas lower-level regions show relatively stronger correspondence with lower layers.
- t-distributed Stochastic Neighbor Embedding Analysis: t-SNE embeds the 4,916 scene representations into two dimensions while preserving similarity relations for visualization.
- t-distributed Stochastic Neighbor Embedding Analysis: Scene and ImageNet images cluster while COCO images scatter more uniformly in scene-processing ROIs, with the strongest Scene clustering in PPA across participants.
- t-distributed Stochastic Neighbor Embedding Analysis: Lower-level visual regions show uniform scattering across datasets, whereas higher-level scene regions show stronger categorical selectivity.This pattern holds across participants.
Usage Notes
BOLD5000 is intended to support joint research across neuroscience and computer vision by enabling more integrated analyses of visual representations.
- The publicly available dataset is designed to create new opportunities for integrated analyses across neuroscience and computer vision.
Discussion
BOLD5000 addresses key gaps in biological-vision datasets by providing substantially greater scale, stimulus diversity, and overlap with computer-vision datasets. The authors acknowledge that its stimulus count and participant count remain limited, motivating larger coordinated studies.
- Discussion: BOLD5000 substantially improves overlap between human neuroimaging data and standard computer-vision datasets.This addresses the field’s identified gap in stimulus overlap alongside gaps in dataset size and diversity.
- Discussion: 5,000 images remain small relative to lifetime human visual experience and the millions of images used to train modern artificial vision systems.The authors propose increasing stimulus images by another order of magnitude in future datasets.
- Discussion: Only four participants were included, reflecting practical constraints on identifying and running suitable participants.The authors argue that detailed individual-level functional data can still be valuable in a small-N design.
- Discussion: Future expansion would require partially overlapping stimulus sets across more participants and methods for stitching the resulting data together.The authors suggest coordinated collaboration among many laboratories for such a scale of human neuroscience.
Data Citations
The cited dataset is BOLD5000, published in 2018.
- Data Citations: BOLD5000 is identified as a 2018 dataset.The passage lists Brain, Object, and Landscape alongside the dataset name.