Source-linked AI summary
Genuine Information Needs of Social Scientists Looking for Data
Andrea Papenmeier, Thomas Krämer, Tanja Friedrich, Daniel Hienert, Dagmar Kern
TL;DR
Social-science data search systems may not accommodate researchers’ complex information needs, which include more than dataset topics. The paper surveys 72 researchers, structures their colleague-style dataset requests into categories, and compares those categories with queries, metadata vocabularies, and search-system capabilities. It finds a mismatch: substantial metadata requirements are absent from query and system support, motivating adaptations to metadata standards and search systems.
Problem
The paper examines whether existing data-search systems, metadata models, and vocabularies cover social scientists’ complex dataset information needs.
Method
An online study collected genuine dataset requests and queries from 72 social scientists, clustered request segments into categories, and compared them with vocabularies and search systems.
Results
Current systems offer only 2.6 “Meta” filters on average, and only one of 30 systems supports specifying meta information directly in query text.
Takeaways & Limitations
Search systems and metadata standards need adaptation to support the broad information needs expressed in social-science dataset requests.
Abstract
from arXiv · showhide
Publishing research data is widely expected to increase its reuse and to inspire new research. In the social sciences, data from surveys, interviews, polls, and statistics are primary resources for research. There is a long tradition to collect and offer research data in data archives and online repositories. Researchers use these systems to identify data relevant to their research. However, especially in data search, users' complex information needs seem to collide with the capabilities of data search systems. The search capabilities, in turn, depend to a high degree upon the metadata schemes used to describe the data. In this research, we conducted an online survey with 72 social science researchers who expressed their individual information needs for research data like they would do when asking a colleague for help. We analyzed these information needs and attributed their different components to the categories: topic, metadata, and intention. We compared these categories and their content to existing metadata models of research data and the search and filter opportunities offered in existing data search systems. We found a mismatch between what users have as a requirement for their data and what is offered on metadata level and search system possibilities.
RELATED WORK
The study builds on research showing that social scientists’ initial information seeking differs from literature searching and remains insufficiently covered by current retrieval models. It investigates genuine dataset needs through colleague-style descriptions, then compares those needs with queries, search systems, and metadata vocabularies.
- Research context: The study focuses on social scientists’ “starting” activity: the initial search for information.Prior work identifies journals, citation tracking, informal channels, and limited library use as recurring characteristics of social scientists’ information seeking.
- Research questions: The study compares genuine information needs with issued queries and evaluates whether existing systems and vocabularies cover those needs.Its research questions address need composition, request-query differences, system coverage, and vocabulary representation.
- Data collection: The researchers asked participants to describe needed quantitative data as if requesting help from a knowledgeable colleague.Participants first reported their research field and topic, then described the data or variables they sought after failing to find a suitable dataset.
- Data collection: Participants wrote dataset requests in their own words, issued a generic-search query, and completed questionnaires about data searching and demographics.The study received ethical clearance from the institution’s ethics committee.
- Analysis: The researchers segmented requests into single characteristics, clustered segments semantically, and re-annotated them using the resulting schema.The final annotation used hierarchical categories at three levels, supporting consistent assignment after iterative clustering.
RESULTS OF DATA ANALYSIS
The analysis organizes dataset-request segments into hierarchical Meta, Topic, and Intention categories. Meta and Topic information dominate the requests, which usually combine multiple category types.
- Clustering: The schema contains three L1 categories—Meta, Topic, and Intention—and 39 L2 categories in total.Meta covers data descriptors such as sample size, geography, and collection method; Topic covers subject matter; Intention captures what researchers want to find out.
- Annotation of Categories: 217 of 450 segments (46%) were Meta, 215 (45%) were Topic, and 44 represented Intention.Socio-demographics was the most common Meta L2 category at 35%, while Politics was the most frequent Topic L2 category at 17%.
- Annotation of Categories: The request syntax was versatile: 35 requests started with Meta information and 34 started with Topic information.No clear ordering pattern emerged for how category information appeared in requests.
- Annotation of Categories: 56% of requests combined Meta and Topic segments, while 22% combined all three L1 categories.Only 12 requests (17%) contained a single L1 category: seven were Meta-only and five were Topic-only.
REQUESTS VS. QUERIES
Dataset requests are significantly longer and more information-rich than search queries, although their broad Topic and Meta proportions are similar. Queries often omit requirements, substitute data properties for socio-demographic details, or compress requests into known-item terms.
- Queries average 4.1 segments versus 6.6 in dataset requests, a significant difference (Wilcoxon, p < .001).Queries also show lower variability: STD = 2.8 versus 4.1 for requests.
- At L1, requests and queries have similar content distributions, with only a minor tendency toward more Topic segments in queries.Requests average 46% Meta, 45% Topic, and 9% Intention segments; queries average 41%, 55%, and 4%, respectively.
- At L2, Data properties rises from 9% of Meta segments in requests to 38% in queries, while Socio-demographics falls from 35% to 18%.
- Only 54% of query segments also appear in dataset requests, reflecting abstraction, elaboration, added information, omitted requirements, and other reformulations.
- Queries can omit Intention requirements, as in reducing a request for representative attitude surveys to “survey, democratic principles.”
- Known-item searches can greatly compress requests, such as representing a detailed longitudinal health-behavior request with “SOEP.”SOEP is the abbreviation for the German Socio-Economic Panel study.
- Vague language appears in 53% of dataset requests but in only one query, indicating a pronounced form difference between natural requests and queries.Vague terms included “different,” “various,” “of the present,” and “as up-to-date as possible.”
COMPARISON WITH EXISTING SYSTEMS
The study evaluated 30 data search systems across five system groups and found that many available facets do not represent researchers’ stated needs. Metadata-aware query interpretation and Intention filtering are especially limited.
- The analysis covered 30 systems spanning five groups: generic engines, generic repositories, social-science engines, social-science repositories, and variable or question search engines.
- Of 183 facets, 95 (52%) matched the study’s information categories, while 88 (48%) addressed information participants did not mention.Examples of non-mentioned facet information include publication year, discipline, contributors, and update date.
- Existing temporal facets usually cover publication or collection years, whereas participants requested forms such as time series, present data, and repeated monthly data.
- Only one system supported filtering by Intention, using the categories Attitudes and Behavior in the UK Data Service Discover system.
- Only Google Dataset Search identified geographical metadata in a query and used the corresponding data field for filtering; other systems matched terms lexically.Lexical matching did not distinguish data from Europe from data about Europe.
SCIENCES
The study compared researchers’ hierarchical categories with three social-science vocabulary systems. Topic coverage was strong in TheSoz and CESSDA, whereas DDI coverage was incomplete or indirect, especially for Meta and Topic categories.
- The vocabulary comparison examined TheSoz, the CESSDA Topic Classification, and the DDI vocabulary family against the study’s emerged categories.These standards were selected because they were developed to describe social-science topics and data.
- TheSoz covered all 14 Topic L3-categories, and CESSDA covered all but the least-mentioned category, events.
- No DDI vocabulary term matched any Topic L3-category directly.
- For Meta, TheSoz matched all but five of 33 L3-categories, while DDI terms matched 19 of 33 categories across five vocabularies.DDI matched only one of ten Socio-demographics L3-categories, and some matches were broader or related rather than exact.
- TheSoz provided explicit counterparts for 16 of the 19 Intention L2-categories, though some concepts were expressed through more specific compound names.
DISCUSSION
The study finds that social scientists’ dataset needs are broad and poorly matched by current queries, metadata vocabularies, and search-system capabilities. It therefore calls for expanded metadata and vocabulary support alongside better search interfaces.
- 72 social scientists provided dataset requests and corresponding search queries for analysis.
- Dataset requests covered broad topic and metadata requirements, with topic information at 45% and meta information at 46%.Intention appeared in 9% of cases.
- Current search systems offered only 2.6 “Meta” filters on average, and just one of 30 systems supported metadata directly in query text.The study concludes that these systems do not suitably support a substantial part of social scientists’ dataset requests.
- Queries were significantly shorter than requests and differed in information content, especially for socio-demographic requirements.Socio-demographics formed 35% of request segments but were not included by participants in their queries.
- Vocabulary coverage was incomplete: DDI matched 16 of 33 “Meta” L3-categories, TheSoz 28, and CESSDA 12.The authors argue that vocabulary standards, documentation tools, metadata records, and search indices all need further support.
CONCLUSION
The study analyzes genuine information needs and search queries from 72 social scientists, comparing their information aspects with existing vocabularies and search systems. It finds that needs extend beyond topic and metadata to intention, while current systems and vocabularies do not adequately cover them.
- The study analyzed dataset requests and queries from 72 social scientists and clustered their information-need aspects.It also compared those aspects with existing vocabularies and search systems.
- Social scientists’ information needs include topic, metadata, and searcher intention requirements.
- Existing search systems lack suitable filters and facets for these information needs.
- The authors conclude that vocabularies and metadata schemes need to be extended with the requested information aspects.