Source-linked AI summary

Gender and the Production of Research Impact

Sanger Wagner, Charles Rahal, Melinda C. Mills

arXiv:2608.26409v1cs.DLstat.AP

TL;DR

The paper addresses limited evidence on which researchers underpin research impact beyond academia and how this production is gendered. It links REF2021 Impact Case Studies with research-output, institutional, and bibliometric data to quantify these patterns. Women comprise 38.16% of named underpinning-research contributors, exceed their representation in output authorships overall and across REF panels, and remain concentrated in less commercial impact domains.

  • Problem

    Which researchers produce the research underpinning documented impact beyond academia, and how is this production shaped by gender?

  • Method

    The study links REF2021 Impact Case Studies with research-output, institutional, and bibliometric records, inferring gender from researchers’ given names.

  • Results

    38.16% of named underpinning-research contributors are women, with higher representation than output authorships overall and in every REF panel; domains also show strong gender sorting.

  • Takeaways & Limitations

    Impact is a heterogeneous form of scientific work: women are better represented in education, health, cultural, and civil-society pathways but remain underrepresented in commercial and industrial pathways.

Abstract

from arXiv · show

Evaluating the impact of scientific research beyond academia --- on policy, health, the economy, and cultural life --- has become a cornerstone of science policy and research-funding allocation worldwide. Yet which researchers produce the research underpinning this impact, and how this production is shaped by gender, remains poorly understood. We combine structured and unstructured records from the United Kingdom's latest Research Excellence Framework, the largest national research assessment currently in operation, with large-scale bibliometric data to quantify gender differences among the researchers underpinning documented impact. Women account for 38.16% of these contributors: underrepresented overall, but with a consistently higher share than in research-output authorships (33.63%), both overall and across all four REF panels. Impact production is also strongly gendered across domains: women are better represented in case studies concerning education, health, cultural, and civil-society impact, whereas those concerning commercialisation pathways such as patenting and manufacturing remain dominated by men. These findings reveal critical inequalities across the pathways that connect research to impact beyond academia, offering crucial evidence for policymakers and academic institutions aiming to build more equitable and representative systems for evaluating scientific contributions.

A Window into Impact: Evidence from a national research evaluation

Using REF2021 Impact Case Studies linked to research-output data, the study shows that women are underrepresented among researchers underpinning documented impact, but their representation varies sharply across disciplines and impact domains. Women are relatively better represented in education, health, cultural, and civil-society pathways, while commercialisation-oriented pathways remain male dominated.

  • A Window into Impact: Evidence from a national research evaluation: REF2021 evaluates research outputs, environment, and impact beyond academia, with impact accounting for 25% of each submission’s overall score.The framework uses Impact Case Studies to document demonstrable benefits beyond academia across four broad disciplinary panels.
  • A Window into Impact: Evidence from a national research evaluation: 17,772 authorships across 6,351 Impact Case Studies identify underpinning researchers, with women comprising 38.16% of gender-inferred authorships.Gender labels were inferred for 16,446 authorships, or 92.5% of the total.
  • A Window into Impact: Evidence from a national research evaluation: Women’s representation varies from majorities in Social Work and Social Policy, English Language and Literature, and Allied Health to minorities in Physics, Engineering, and Computer Science.The cited fields range from 57.30% women in Social Work and Social Policy to 14.05% in Physics.
  • Impact is not a Single Activity: Strong gender sorting across impact domains: Women account for nearly half of researchers in Charity, School, Museum, and NHS case studies, but only around one in five in Patent, Manufacturing, and Startup domains.The reported shares are 47.54% in Charity and 17.89% in Patent, illustrating strong gender sorting across impact domains.
  • Domains, Disciplines, and Institutions in Impact Production: Panel B is associated with a 27.2-percentage-point lower share of women, attenuated to 20.0 percentage points after accounting for impact domains.The remaining association indicates that disciplinary context and impact-domain composition both relate to gender representation.
  • Implications for Understanding Gender and Research Impact in Science: Women are more strongly represented among ICS underpinning-research authorships than output authorships overall and in every REF panel, with the pattern holding in 21 of 34 fields.The estimates concern researchers credited with conducting underpinning research, not every person who created or delivered the impact.
  • Impact as Opportunity and Constraint: Impact can provide pathways beyond publication-based metrics while also concentrating women in less commercial forms of engagement.The paper argues that treating impact as one category obscures the gendered labour supporting distinct economic and societal pathways.

S1 Data Sources

The study combines public REF2021 case-study and output records with institutional data and bibliometric enrichment. These sources support analyses of credited underpinning researchers, publication-output benchmarks, and institutional characteristics.

  • S1 Data Sources: The public REF2021 Impact Case Study database provides 6,361 records and structured fields including institution, UoA, panel, and impact descriptions.These records are used to analyse gender representation among researchers credited with underpinning documented impact.
  • S1 Data Sources: The published REF2021 research-output workbook supplies submitted outputs and identifiers, while Dimensions provides author metadata for gender analysis.DOI and ISBN matches connect REF outputs to bibliographic author records.
  • S1 Data Sources: Institutional measures include submitted-staff FTE, doctoral degrees awarded, total research income, and institution-type indicators.Institution types include Oxbridge, Russell Group, Red Brick, and Ancient universities.
  • S1 Data Sources: The analysis uses Dimensions API queries in DOI and ISBN batches, caches returned records, and prefers DOI matches when both identifiers are available.This procedure supports reproducible bibliographic enrichment and record linkage.

S2 Data Extraction and Variable Construction

The study constructs researcher, gender, and impact-domain variables from REF2021 case-study records using PDF extraction, name-based gender inference, and conservative multi-label classification. Classification robustness is assessed against manual review, regular-expression rules, and alternative language-model assignments.

  • Name-based gender inference: Gender labels were inferred from candidate given names using two offline lexicons in fixed precedence order, with initials, unresolved particles, and ambiguous outputs assigned unknown.The lexicons are collapsed into men, women, and unknown categories before person-level labels are re-aggregated to cases.
  • Impact-domain construction: Impact domains were classified from normalized text across five structured REF fields, using conservative multi-label rules that permit multiple domains but code uncertain cases false.Labels are positive only when a domain is materially involved in the claimed impact; passing mentions and background context are excluded.
  • Impact-domain construction: The classifier distinguishes impact routes including charity, startups, patents, museums, NHS, drug trials, schools, legislation, heritage, manufacturing, and software.Each domain has a material-involvement rule, such as patents or licensing for patent impact and industrial production outcomes for manufacturing impact.
  • Validation and cross-checks: 90.1% agreement was observed between regular-expression and GPT-5.5 classifications, while pairwise agreement among language-model variants ranged from 93.6% to 96.0%.Agreement was typically highest for narrower domains and disagreement concentrated in broader domains describing indirect impact routes.

S3 Regression Models and Robustness

The study models women’s representation among impact-case-study contributors while progressively controlling for disciplinary panels, institutional characteristics, and impact domains. Robustness analyses show that domain-level gender associations persist across estimators, finer disciplinary controls, and alternative domain classifications.

  • S3.1 Regression models: Weighted OLS estimates women’s case-study share while progressively accounting for REF panels, institution types, and 11 impact domains.The unit is a case study with at least one gender-identifiable contributor, weighted by the number of such contributors; binomial GLMs provide a robustness check.
  • S3.2 Domain composition and gender representation: physics vs chemistry: Physics has 14.05% women among impact authorships versus 6.94% among output authorships, while Chemistry has 19.11% versus 24.97%, respectively.Physics therefore has a positive impact–output gap, whereas Chemistry has a negative one.
  • S3.2 Domain composition and gender representation: physics vs chemistry: Physics contains more case studies in higher-representation domains, including School, Museum, and Charity, whereas Chemistry is more concentrated in lower-representation domains.The comparison uses domain shares because Physics has 169 case studies and Chemistry 113.
  • S3.3 Alternative model specification: Panel B remains the strongest negative association after institutional and domain controls, with a GLM coefficient of −1.004 and an OLS estimate of −20.0 percentage points.The two estimates use different scales; the robustness criterion is preservation of signs, ordering, and substantive interpretation rather than numerical equality.
  • S3.4 Robustness to finer discipline controls: Replacing REF panels with finer Unit of Assessment controls leaves Charity, Museum, NHS, and School positively associated, while Patent, Drug Trial, Manufacturing, Software, and Startup remain negative.The domain associations are therefore not simply proxies for broad panel composition or a few highly gendered disciplines.
  • S3.5 Robustness to impact-domain classification: Across alternative regular-expression and LLM-based classifications, commercial domains remain lower in women’s representation, while schools, museums, charities, NHS, and heritage remain comparatively more balanced.Panel B’s lower representation is also preserved across classifications.

S4 Qualitative Interviews

The study uses interviews with REF2021 impact-assessment participants to contextualize quantitative patterns in the case-study data. These interviews support interpretation rather than constituting an independent analysis.

  • S4 Qualitative Interviews: Interviews conducted between March and December 2023 covered REF2021 assessment-panel members and other participants directly involved in impact assessment.They originated in prior research supporting the British Academy report The SHAPE of Research Impact.
  • S4 Qualitative Interviews: The interviews are used as supporting qualitative evidence to inform interpretation of patterns observed in REF impact case-study data.They are not presented as an independent analysis.

S5 Supplementary Descriptive Results

Supplementary descriptive results show substantial gender variation across REF panels, Units of Assessment, and impact domains, while text-based analysis provides an additional check on these patterns.

  • S5.1 Descriptive statistics by REF main panel: 45.51% of gender-identifiable impact authorships in Panel A were women, while panel-level aggregates conceal substantial disciplinary heterogeneity.Panel C is largest by submitted staff and impact case studies, whereas Panel A leads in research income and submitted outputs.
  • S5.2 Descriptive statistics and representation by unit of assessment: 14.05% of impact authorships in Physics and 19.11% in Chemistry were women, illustrating contrasting impact–output gaps across two male-dominated STEM fields.Physics had a positive impact–output gap, whereas Chemistry had a negative one.
  • S5.3 Descriptive statistics and gender representation across impact domains: 47.54% of authors in Charity case studies were women, compared with 17.89% in Patent case studies, showing sharply different representation across impact domains.Women’s representation was also high in School, Museum, and NHS domains and low in Manufacturing, Startup, Software, and Drug Trial.
  • S5.3 Descriptive statistics and gender representation across impact domains: Word-level analysis tested whether terms in impact narratives were associated with higher or lower shares of women contributors.The analysis used substantive narrative sections and vocabulary terms appearing often enough to support comparison.

S6 Abbreviations and Figure Labels

The supplementary materials define abbreviated REF panels and map shortened figure labels to official Units of Assessment.

  • Abbreviations: Panels A–D denote medicine, health and life sciences; physical sciences, engineering and mathematics; social sciences; and arts and humanities, respectively.The full REF2021 Unit of Assessment names are too long for figure axes.
  • Figure labels: Table S2 maps shortened Unit of Assessment labels used in figures to their official REF2021 names.Institution-type and impact-domain labels appear in full in all figures.

Supplementary Figures

The supplementary figures assess classification agreement, text associations, estimator robustness, disciplinary controls, and the stability of gender patterns across alternative domain-classification approaches.

  • Figure S1: Figure S1 compares thematic breadth, topic-level agreement, aggregate method comparisons, and disagreement rates across LLM and regular-expression classifiers.Thematic breadth is shown through the empirical cumulative distribution of domains assigned per case study.
  • Figure S2: Figure S2 relates words in impact case studies to higher or lower shares of women authors, with false-discovery-rate significance shown for prevalence differences.It provides a text-based check on domain-classification results.
  • Figure S3: Figure S3 reproduces the Figure 2b model using a binomial GLM with a logit estimator as an alternative estimation strategy.The GLM uses a binomial likelihood and log-odds scale, complementing the main OLS coefficient plot.
  • Figure S4: Figure S4 replaces REF panels with finer Units of Assessment in Model Three, estimating the specification with both OLS and a related generalised linear model.The figure tests robustness to more granular disciplinary controls.
  • Figure S5: Figure S5 compares domain-level gender patterns using regular expressions and three alternative large language modelling approaches.The alternative approaches are equivalent to the analysis shown in Figure 2a.

Supplementary Tables

The supplementary tables provide abbreviation mappings, descriptive statistics, domain distributions, regression results, and panel- and field-level classifications supporting the paper’s analyses.

  • Reference tables: Table S1 lists abbreviations, while Table S2 maps shortened Unit of Assessment labels to official REF2021 names.These tables support interpretation of the main text and figures.
  • Impact-domain descriptives: Tables S3 and S4 report impact-case-study distributions and women’s representation overall and by REF main panel across impact domains.Because domain assignment is multi-label, case studies may contribute to multiple domains.
  • Unit-of-Assessment descriptives: Table S5 reports descriptive statistics and women’s representation in impact case studies and research outputs across REF Units of Assessment.It provides the UoA-level basis for disciplinary comparisons.
  • Regression and field comparisons: Table S6 compares Physics and Chemistry domain distributions and provides associated OLS and GLM estimates from Model Three.Table S7 gives the full regression results for models of women’s representation in impact case studies.
  • Statistical notation: Standard errors are reported in parentheses, with significance markers for p < 0.10, p < 0.05, and p < 0.01.The notation distinguishes three conventional significance thresholds.
  • Panel-specific classifications: Table S8 reports Panel D case studies classified as Patent and Drug Trial by Unit of Assessment using LLM classification.Table S9 reports descriptive statistics and women’s representation by REF main panel.
Loading 2608.26409v1…