Source-linked AI summary
The Diversity-Innovation Paradox in Science
Bas Hofstra, Vivek V. Kulkarni, Sebastian Munoz-Najar Galvez, Bryan He, Dan Jurafsky, Daniel A. McFarland
TL;DR
The paper asks whether the diversity-innovation paradox extends to science and examines this question using doctoral-career and publication data. It finds that underrepresented groups produce more novelty, but their contributions receive less uptake and weaker career returns.
Problem
The paper examines whether diversity's association with innovation and career success produces a paradox for underrepresented scientists.
Method
The study analyzes 1,208,246 dissertations and uses text-based concept-link measures to identify meaningful scientific innovations.
Results
Underrepresented groups produce higher novelty rates, but their novel contributions receive less uptake and yield lesser returns to scientific careers.
Takeaways & Limitations
The findings identify a diversity-innovation paradox in science: underrepresented groups' innovations are not equally valued or rewarded.
Takeaways & Limitations
Analyzing full text involves theoretical difficulties, including deciding how closely concepts must co-occur to form a meaningful link.
Abstract
from arXiv · showhide
Prior work finds a diversity paradox: diversity breeds innovation, and yet, underrepresented groups that diversify organizations have less successful careers within them. Does the diversity paradox hold for scientists as well? We study this by utilizing a near-population of ~1.2 million US doctoral recipients from 1977-2015 and following their careers into publishing and faculty positions. We use text analysis and machine learning to answer a series of questions: How do we detect scientific innovations? Are underrepresented groups more likely to generate scientific innovations? And are the innovations of underrepresented groups adopted and rewarded? Our analyses show that underrepresented groups produce higher rates of scientific novelty. However, their novel contributions are devalued and discounted: e.g., novel contributions by gender and racial minorities are taken up by other scholars at lower rates than novel contributions by gender and racial majorities, and equally impactful contributions of gender and racial minorities are less likely to result in successful scientific careers than for majority groups. These results suggest there may be unwarranted reproduction of stratification in academic careers that discounts diversity's role in innovation and partly explains the underrepresentation of some groups in academia.
Introduction · Innovation as Novelty and Impactful Novelty in Text
The paper frames a diversity-innovation paradox: underrepresented scholars are expected to innovate more and build successful careers, yet their contributions may be discounted. Using near-population dissertation and publication data, it measures novelty, uptake, and career returns across demographic groups.
- Introduction: Diversity is expected to generate more scientific innovation and successful careers, but persistent career inequalities create a diversity-innovation paradox.The proposed explanation is that innovations produced by some groups are discounted, reducing their scientific impact and career success.
- Introduction: The study identifies the diversity-innovation paradox in science and explains why it arises through a system-level analysis.It compares demographic groups’ novelty rates, uptake by other scholars, and effects on successful research careers.
- Introduction: The analysis spans three decades, all scientific disciplines, and all US doctorate-awarding institutions.This scope enables comparisons of minority and majority scholars’ scientific novelty, uptake, and career outcomes.
- Innovation as Novelty and Impactful Novelty in Text: The dataset contains nearly all US PhD theses and metadata from 1977-2015, including names, advisors, institutions, titles, abstracts, and disciplines.Records are linked to Census and Social Security data for demographic inference and Web of Science publications for research-career tracking.
- Innovation as Novelty and Impactful Novelty in Text: The study measures novelty as the number of new conceptual links introduced in theses using phrase extraction and structural topic modeling.It identifies meaningful concept pairs and sums their novel co-occurrences within each thesis.
- Innovation as Novelty and Impactful Novelty in Text: Impactful novelty is measured as uptake per new link, or the average future adoption of a thesis’s unique conceptual recombinations.Novelty had Mean = 9.026 and 20.9% of students introduced no links; uptake per new link had Mean = .790.
- Innovation as Novelty and Impactful Novelty in Text: Novelty does not automatically constitute innovation, and adoption is not required for innovation because structural processes may shape which novelty gets adopted.Conceptual recombination also avoids discipline-specific journal-indexing priorities and the varied reasons scholars cite other work.
Results
Underrepresented scientists introduce more novelty, but their contributions receive less uptake and weaker career rewards than those of majority groups. This discounting is partly associated with the greater semantic distance of their novel conceptual linkages.
- Who introduces novelty and whose novelty is impactful?: Students from underrepresented genders or races introduce more novel conceptual linkages, while similar-gender representation increases uptake by others.Underrepresented genders (p < .001) or races (p < .05) predict more new links, whereas gender representation predicts greater uptake (p < .01).
- Who introduces novelty and whose novelty is impactful?: Women and non-white scholars introduce more novelty but have less impactful novelty than men and white scholars.Both differences are statistically significant: novelty (p < .001) and impactful novelty (p < .05).
- Why is novelty less impactful?: Underrepresented genders introduce slightly more semantically distant linkages, and distal new links receive far less uptake.Gender-underrepresentation and women’s greater distal novelty are both reported at p < .001; distal novelty is inversely related to impactful novelty (p < .001).
- Novelty and scientific careers: Conceptual novelty and impactful novelty both increase the likelihood of becoming research faculty or continuing as a researcher (all p < .001).These career models hold institution, discipline, and graduation year constant.
- Novelty and scientific careers: Numerically underrepresented genders have approximately 5% lower odds of becoming research faculty and 6% lower odds of sustaining research careers than gender majorities.Numerically underrepresented races have 25% lower odds of becoming research faculty and 10% lower odds of continuing research endeavors than majorities; all p < .001.
Discussion … Concept extraction from scientific text
The study finds that demographic minorities generate more scientific novelty, but their contributions receive less uptake and weaker career rewards. It measures these patterns using a large dissertation corpus and structural topic models that identify distinctive concepts from topical vocabularies.
- Discussion: Underrepresented groups produce higher rates of scientific novelty, yet their novel contributions are adopted less often and yield weaker career returns than majority groups.The discounting applies to both uptake and successful scientific careers, including highly impactful novelty.
- Discussion: Novel conceptual linkages from women and non-white scholars receive less uptake than comparable linkages from dominant or majority groups.For gender minorities, lower uptake is partly explained by the greater distance of their novel conceptual connections from prevailing field conversations.
- Discussion: Minorities must innovate at higher levels to achieve similar career success, while their careers may end prematurely despite generating novel discoveries.The paper argues that this stratification warrants continued evaluation of bias in faculty hiring, research evaluation, and publication practices.
- Data: 1,208,246 dissertations from 1977 to 2015 cover approximately 86% of U.S. doctorates across all disciplines over three decades.The dataset includes candidate, award-year, university, thesis-abstract, and advisor metadata.
- Data: The study follows doctoral recipients’ subsequent academic and research careers using their earliest intellectual footprints.The discussion describes following over a million U.S. students’ careers and publishing or faculty outcomes.
- Concept extraction from scientific text: Innovation is operationalized as combining relevant terms from topical lexicons rather than merely combining function words or arbitrary vocabulary.Structural Topic Models identify latent themes, and FREX scores select terms that balance frequency with exclusivity.
- Concept extraction from scientific text: Inferential analyses use theses from 1982 to 2010, allowing five years for concept-space accumulation and five years for newly graduated students’ novelty uptake.Year fixed effects further address left- and right-censoring by comparing students within rather than across years.
Outcome variable–Novelty and impactful novelty · Outcome variable–Distal novelty
The paper measures thesis novelty by identifying statistically meaningful new concept links and their subsequent uptake, then quantifies distal novelty through semantic distance between linked concepts. These measures capture both the production and impact of scientific novelty, including cross-field conceptual connections.
- Outcome variable–Novelty and impactful novelty: Novelty counts the number of statistically meaningful concept links a student’s thesis introduces, after filtering topic-model terms and spurious links.Links introduced concurrently by students in the same year are counted for both students.
- Outcome variable–Novelty and impactful novelty: Mean = 9.026 new links per thesis, while 20.9% of students introduce no new links.The distribution has Median = 4 and SD = 13.744.
- Outcome variable–Novelty and impactful novelty: Impactful novelty measures the uptake of a thesis’s new links in subsequent theses, normalized by the number of new links.Uptake per new link has Mean = .790, Median = .333, and SD = 3.079.
- Outcome variable–Novelty and impactful novelty: Both novelty measures positively correlate with publication productivity and citations among students who publish.The correlations are reported across the different K and FREX scenarios.
- Outcome variable–Distal novelty: Distal links connect concepts from distinct co-occurrence clusters, whereas proximal links connect concepts within the same semantic cluster.For example, genetic_algorithm–hiv-1 is distal, while fracture_behavior–ceramic_composition is proximal.
- Outcome variable–Distal novelty: Expert coders show average Cohen’s Kappa = .46, and their assignments predict ~95% of true distal links with distance scores above .8.Distal links often connect concepts from different fields or creative metaphors, while 15-20% are difficult to interpret substantively.
Outcome variable–Careers · Main covariates
Career outcomes are measured using both research-faculty placement and continued research activity, while gender and race are inferred from names and modeled through several representation measures. These proxies broaden career coverage but simplify identity and exclude unclassified cases, with main conclusions remaining robust in restricted analyses.
- Outcome variable–Careers: Research faculty is a conservative career-success proxy: becoming a primary advisor of PhD students at a US PhD-granting university.The outcome has mean = .066 and captures transition from student to mentor with a student lineage.
- Outcome variable–Careers: Continued research is a broader proxy, defined as publishing academically at least once within five years after the PhD or becoming research faculty.Continued research has mean = 0.319 and includes institutions such as liberal arts colleges, think tanks, industry, and international destinations.
- Outcome variable–Careers: The ProQuest–Web of Science linkage follows students’ subsequent publication and career outcomes using ~38 million records from 1900 to 2017.The linked data capture ensuing careers and research output.
- Main covariates: Using a training set of N = 20,264 private-university scholars, a threshold algorithm estimates gender and race from names.Race is classified as White, Asian, or Underrepresented minorities; gender is assigned using the resulting and supplementary methods.
- Main covariates: Main covariates include same-gender and same-race representation, minority status within discipline-years, and white/non-white status for modeling innovation and careers.Mean = .576 for % Same-gender, Mean = .625 for % Same-race, Gender minority mean = .336, and Racial minority mean = .246.
- Main covariates: The identity measures are coarse proxies: they omit unclassified gender and race cases and cannot fully represent self-identification or finer-grained racial categories.Names may better capture how individuals are perceived by others than their self-identification.
- Main covariates: Main substantive conclusions remain robust when analyses restrict students to names overwhelmingly associated with one race.Unknown genders and races do not qualitatively change the distinctions shown in Figure 2C and F.
Confounding Factors … Diversity and Innovation
The study motivates a diversity-innovation mechanism in which underrepresented scientists bring distinct perspectives that may generate novel connections, and uses controlled statistical models to examine these outcomes. Supplementary materials document the study’s data, methods, funding, disclosures, and availability.
- Analytical strategy: Scientific novelty and impactful novelty are right-skewed event counts or rates analyzed with negative binomial regression, while distal novelty uses linear regression.Career outcomes—becoming and sustaining research faculty—are binary and analyzed with logistic regression.
- Analytical strategy: Institution, academic discipline, and graduation year are held constant to isolate predictors from university, field, and cohort confounding factors.The models also use doctorate-weighted data to account for university selectivity and improve generalizability to the US scholarly population.
- Funding: The study received support from two National Science Foundation grants and a grant from the Dutch Organization for Scientific Research.The authors also declare no competing interests.
- Data and materials availability: Data access followed Stanford-approved protocol 12996, with ProQuest permission for dissertation analysis and availability of dissertation, Web of Science, and replication resources.The replication code and top terms from the K = 500 structural topic model are available through the stated repositories or providers.
- Supplementary Text: The supplementary text covers innovation measurement, concept extraction, meaningful-link identification, student demographics, academic discipline, population coverage, and inferential models.These materials are listed as components of the supplementary information.
- Diversity and Innovation: Underrepresented groups may generate novel connections by drawing relationships between ideas that majority-group members missed or ignored.Their distinct experiences and perspectives can broaden the variety of viewpoints brought to scientific research.
- Diversity and Innovation: Gender and racial diversity can create heterogeneous pools of thinkers that improve innovation, problem solving, and prediction accuracy.The passage attributes these benefits to diversity bonuses arising from heterogeneity in collective thought.
Measuring Innovation Through Citations, Keywords, and Text · Structural Topic Models for Concept Extraction · Structural Topic Models
The study measures scientific innovation through novel recombinations of concepts extracted from dissertation text, extending citation- and keyword-based approaches. Structural Topic Models identify latent research topics and provide robust concept representations used to operationalize novelty.
- Measuring Innovation Through Citations, Keywords, and Text: Innovation is measured through novel recombinations of concepts in scientific text, extending prior approaches based on citations, keywords, and bibliographic sources.Text-based recombinations preserve the explicit meaning of concept combinations while addressing limitations of citation and keyword measures.
- Measuring Innovation Through Citations, Keywords, and Text: A limitation is that concept introductions identified in ProQuest dissertations may have appeared earlier in peer-reviewed journals or other corpora.The dissertation corpus nevertheless indicates which dissertations are novel relative to others and which students produce distinctive early knowledge.
- Measuring Innovation Through Citations, Keywords, and Text: The dissertation corpus captures early innovative sparks across academic fields, including slower, book-oriented and faster, proceedings-oriented sciences.Language metrics are less affected by indexing differences, journal guidelines, and citation practices than citation records.
- Structural Topic Models: Structural Topic Models are fit to approximately 1.2 million dissertation abstracts, modeling topic prevalence as a linear function of doctorate year.STMs represent documents as mixtures of latent thematic dimensions and learn topic-word and document-topic distributions through variational expectation-maximization.
- Structural Topic Models: The model treats scientific innovation as novel combinations of terms that are distinct and heavily used within latent research areas.Compared with top TF-IDF terms, STM-based extraction identifies terms with significant roles in underlying thematic dimensions.
- Structural Topic Models: Internal and external validation identify approximately K = 400-600 topics as optimal, with the main analysis using K = 500.The K = 500 STM supplies vocabulary weights for extracting concepts describing latent dimensions in the corpus.
- Structural Topic Models: Results remain robust across alternative concept-extraction specifications and topic counts of 400, 500, and 600.The alternatives vary emphasis on frequency, exclusivity, or a balance between them, with concepts drawn from highest-FREX-score terms.
Preprocessing Texts and Fitting STMs · Internal validation · External validation
The analysis preprocesses dissertation text, fits Structural Topic Models across K = 50-1000, and validates topic dimensionality using internal coherence/exclusivity and external agreement with author-provided fields and keywords. Both validation approaches indicate an optimal solution around K = 400-500, with external similarity declining beyond K = 500.
- Preprocessing Texts and Fitting STMs: The preprocessing removes stand-alone numbers, punctuation, English stop words, and special characters while retaining numbers embedded in substantive terms such as H2S.The procedure also stems words with Snowball and removes tokens appearing only once across documents.
- Preprocessing Texts and Fitting STMs: The analysis stems tokens, removes singleton terms, extracts unusually frequent n-grams, and fits STMs from K = 50-1000 in incremental steps.Models are trained for 20 epochs, with steps of 100 when K > 600 to reduce computing time.
- Internal validation: Internal validation selects topic dimensionality by jointly evaluating semantic coherence and exclusivity across different values of K.Coherence captures co-occurrence among high-probability words, whereas exclusivity captures their distinctiveness across topics.
- Internal validation: Because coherence can be trivially increased with fewer topics, the analysis seeks a K where coherence and exclusivity plateau rather than relying on coherence alone.Common words can appear highly coherent while having low exclusivity because they co-occur across many topics.
- Internal validation: Internal validation places the likely topic limit in the range K = 400-.This range is identified from the behavior of coherence and exclusivity in Figures S1-A and S1-B.
- External validation: External validation compares STM-based document distances with distances based on dissertation authors’ academic fields and keywords.It uses cosine similarity among topic-mixture vectors for a constant random sample of 1000 documents and defines related pairs using the sample median.
- External validation: The external comparison represents STM and bigram relations as two document networks and measures dyadic overlap with the Matthew correlation coefficient.The coefficient accounts for true negatives as well as true and false positives and negatives.
Consistency · The “Right” K · Concept Extraction With FREX
The analysis identifies a stable topic range around K = 400–600, uses K = 500 as the main specification, and finds qualitatively similar results across nearby K values. Concept extraction balances term frequency and exclusivity through FREX, with sensitivity analyses robust across nine specifications.
- Consistency: Fowlkes-Mallows consistency rises with K and stabilizes between K = 400 and K = 600, indicating a plateau in topic-assignment overlap.Raw FM scores suggest more than two-thirds of document-to-topic assignments are stable from K = 400.
- The “Right” K: The analysis uses K = 500 in the main text because the consistency metrics appear to plateau around that value.The authors reject the idea of a uniquely “right” K because it would imply complete knowledge of the topic universe.
- The “Right” K: Choosing K = 400, K = 500, or K = 600 leaves the key results mostly qualitatively similar.Concept/link introduction and uptake are measured similarly across these choices using low, medium, and high FREX weights.
- The “Right” K: The plausible “right” K for scientific topics at the specialization-within-disciplines level likely lies between 400-600.This range is based on the observed metric plateau rather than a claim that one K is uniquely correct.
- Concept Extraction With FREX: FREX extracts concepts that balance how common and distinctive terms are within topics, avoiding terms that are merely generic or idiosyncratic.FREX compounds weighted frequency and exclusivity to identify terms informative for concepts.
- Concept Extraction With FREX: The extraction tests three FREX weighting schemes—50/50, 75/25, and 25/75—across K = [400-600] in 100-topic increments, yielding nine scenarios.The analysis extracts the top-500 FREX-words per topic and calculates innovation variables for each scenario.
- Concept Extraction With FREX: Sensitivity analyses produce robust results for novelty, impactful novelty, and recognition across the nine scenarios.The main text reports the scenario equally balancing frequency and exclusivity at K = 500.
On Analyzing Abstracts Versus Full Texts · The PMI Score to Identify Meaningful Links · Student Gender and Race
The study analyzes dissertation abstracts as scalable, substantively meaningful approximations of full texts, identifies meaningful concept links using PMI, and infers student gender and race from names. Name-based classifications are validated against self-reported data, supplemented with an external method, and tested for robustness using higher-precision cases.
- On Analyzing Abstracts Versus Full Texts: Abstracts approximate full-text knowledge while avoiding full-text access, sampling, scalability, and computational constraints.Abstracts are easier to obtain and typically require fewer computational resources; full-text co-occurrence also raises unresolved distance choices.
- On Analyzing Abstracts Versus Full Texts: Abstract co-occurrences are treated as substantively meaningful because abstracts summarize key contributions in ~10 sentences.The analysis assumes numerical minorities do not retain innovations in abstracts at systematically different rates than majorities.
- The PMI Score to Identify Meaningful Links: PMI scores quantify whether a concept link occurs more often than chance using its joint and marginal probabilities.For link L = (a, b), the score uses Pr(a, b), Pr(a), and Pr(b).
- The PMI Score to Identify Meaningful Links: The analysis filters spurious recombinations with a rank-based PMI cutoff, requiring terms to occur in at least 10 theses and considering the top 10 million links.This preserves opportunities for novelty while removing obviously meaningless links.
- Student Gender and Race: Student gender and race are inferred from first and last names because the ProQuest corpus lacks these records.Race probabilities come from 2000 and 2010 US Census data, while gender probabilities come from Social Security Administration data.
- Student Gender and Race: The name-matching validation covered 83.9% of self-reported race cases and 87.5% of self-reported gender cases.The matching datasets contained Ntotal = 24,150 and Nmatch = 20,264 for race, and Ntotal = 35,469 and Nmatch = 31,026 for gender.
- Student Gender and Race: The inferred classifications achieved correlations of .91 for gender and .83, .93, .73, and .25 for white, Asian, Hispanic, and African and Native American race categories, respectively.Correct-identification rates were 97.2% for white, 93.4% for Asian, 70.4% for Hispanic, and 9.9% for African and Native American cases.
- Student Gender and Race: A supplementary full-name method improves Hispanic and African American classification, while higher-precision race analyses produce qualitatively similar results.Unknown cases comprise ~8.5% for gender and ~10.8% for race; supplementary-method precision is .83 for Hispanic and .74 for African American names.
Academic Discipline · Population Coverage and Data Weights · Inferential Models
The study infers missing academic disciplines with a machine-learning classifier, draws on approximately 86% of US doctorates, and applies institution-year survey weights. Inferential models estimate novelty, impactful novelty, and career success while adjusting for key confounders.
- Academic Discipline: Departments are mapped to National Research Council categories using fuzzy string matching, manual verification, and a successfully matched ground truth of N = 178,511 dissertations.Unmapped departments are inferred using the classifier trained on dissertations with identified NRC departments.
- Academic Discipline: A Random Forest Classifier infers dissertation departments with 96% precision (NDISCIPLINE = 84).It uses subject categories, keywords, abstract topics, Word2Vec representations, and degree-granting university features.
- Population Coverage and Data Weights: Approximately 1.2 million doctorates were awarded in the United States during 1977-2015, and ProQuest covers approximately 86% of them.ProQuest and national doctorate trends are highly similar over time.
- Population Coverage and Data Weights: For inferential analyses, 1982-2010 observations are weighted by institution-year doctorate totals to account for selectivity in thesis filing.Weights compare ProQuest institution-year doctorate shares with National Science Foundation census shares.
- Inferential Models: Equation (2) models expected counts or rates of link introductions and uptake per new link, while equation (3) models average link distality.Uptake per new link represents impactful novelty and is modeled as a rate rather than an integer event count.
- Inferential Models: Equation (4) models career success as becoming a faculty researcher or sustaining a research career.All equations apply to individual student j.
- Inferential Models: Models include gender and race representation as main predictors and institution, discipline, and year as confounding factors.Coefficients represent covariate associations with the modeled outcome Y.
- Inferential Models: A logged number of new links is included as an offset so uptake per new link can be modeled with rate increases or decreases.The offset converts expected counts into rates by dividing counts by exposure.
Linking ProQuest to Web of Science · Supplementary References
The study links ProQuest dissertations to Web of Science authors across pre- and post-2009 publication corpora using conservative, multi-criteria matching. Supplementary analyses assess linkage precision, novelty measurement, topic modeling, representation, and associations between novelty and later publication outcomes.
- Linking ProQuest to Web of Science: The linkage combines pre-2009 and post-2009 Web of Science author-disambiguation datasets to improve accuracy and coverage across the full time range.The pre-2009 data use author clusters, whereas post-2009 data contain Clarivate-disambiguated authors.
- Linking ProQuest to Web of Science: The conservative linking rules prioritize reducing mistakenly linked clusters over recovering all undiscovered linkages.Strict requirements include a 75% article-subset match or at least one matching email address, followed by manual and automated verification on random samples.
- Linking ProQuest to Web of Science: 97% precision was inferred for the alignment between the pre-2009 and post-2009 author-cluster datasets.The estimate came from online, self-labeled publications provided by Clarivate Analytics.
- Linking ProQuest to Web of Science: ProQuest-to-WoS matching uses successive criteria based on advisor or advisee co-authorship, institutions, keywords, names, and dissertation–publication text similarity.The procedure evaluates highest-confidence criteria first and uses maximum bipartite matching so each author is linked to at most one author on the other side.
- Linking ProQuest to Web of Science: The matching procedure reduces mismatches by requiring publication timing consistent with graduation and additional evidence beyond name and textual similarity.It excludes publication histories centered 15 years after or 10 years before graduation, generally avoids predominantly pre-graduation publishing, and rejects weakly supported matches.
- Supplementary References: Supplementary analyses report stable novelty patterns from approximately 1982, robustness across topic-model settings, and positive relationships between novelty, publications, and accumulated citations.Sensitivity analyses vary K and FREX weighting; regression models include discipline, PhD graduation year, and PhD university fixed effects.