Source-linked AI summary

Defining and identifying Sleeping Beauties in science

Qing Ke, Emilio Ferrara, Filippo Radicchi, Alessandro Flammini

arXiv:1505.06454v1physics.soc-phcs.DLcs.SI

TL;DR

Previous Sleeping Beauty studies relied on arbitrary thresholds and limited bibliographic datasets, leaving the prevalence of delayed recognition uncertain. This paper introduces a parameter-free measure and applies it to large-scale, multidisciplinary citation data, finding that Sleeping Beauties are not exceptional but lie on a continuous spectrum of delayed recognition. The results also show delayed recognition across disciplines and caution against short-term citation metrics for scientific impact.

  • Problem

    Previous methods used arbitrary sleeping-time and citation thresholds on small or monodisciplinary datasets, limiting evidence about the prevalence of Sleeping Beauties.

  • Method

    The study introduces a parameter-free measure of Sleeping Beauty strength and applies it in a systematic large-scale analysis.

  • Results

    Papers with long dormant periods followed by rapid citation growth are not exceptional outliers but extreme cases within continuous, heterogeneous distributions.

  • Takeaways & Limitations

    Delayed recognition is a continuous citation-dynamics phenomenon whose prevalence and cross-disciplinary forms are revealed by broad bibliographic analysis.

  • Takeaways & Limitations

    The citation-pattern analysis has a limitation acknowledged by the authors.

Abstract

from arXiv · show

A Sleeping Beauty (SB) in science refers to a paper whose importance is not recognized for several years after publication. Its citation history exhibits a long hibernation period followed by a sudden spike of popularity. Previous studies suggest a relative scarcity of SBs. The reliability of this conclusion is, however, heavily dependent on identification methods based on arbitrary threshold parameters for sleeping time and number of citations, applied to small or monodisciplinary bibliographic datasets. Here we present a systematic, large-scale, and multidisciplinary analysis of the SB phenomenon in science. We introduce a parameter-free measure that quantifies the extent to which a specific paper can be considered an SB. We apply our method to 22 million scientific papers published in all disciplines of natural and social sciences over a time span longer than a century. Our results reveal that the SB phenomenon is not exceptional. There is a continuous spectrum of delayed recognition where both the hibernation period and the awakening intensity are taken into account. Although many cases of SBs can be identified by looking at monodisciplinary bibliographic data, the SB phenomenon becomes much more apparent with the analysis of multidisciplinary datasets, where we can observe many examples of papers achieving delayed yet exceptional importance in disciplines different from those where they were originally published. Our analysis emphasizes a complex feature of citation dynamics that so far has received little attention, and also provides empirical evidence against the use of short-term citation metrics in the quantification of scientific impact.

I. MATERIALS

The beauty coefficient B compares a paper’s citation history with a publication-year-defined reference line, incorporating delayed and intense citation growth. The method uses the full citation trajectory and avoids arbitrary thresholds for sleeping time or awakening intensity.

  • Beauty coefficient: B compares a paper’s citation history with a reference line determined by its publication year and citation maximum.The reference line connects the citation counts at publication and at the year of maximum yearly citations.
  • Beauty coefficient: The reference line ℓ_t connects (0, c_0) and (t_m, c_t_m) in the time-citation plane.Here, t denotes the paper’s age in years after publication.
  • Beauty coefficient: B sums the ratios between the reference-line value ℓ_t minus actual citations c_t and max{1, c_t} over 0 ≤ t ≤ t_m.This construction compares the citation trajectory with the reference line throughout the interval up to the maximum-citation year.
  • Beauty coefficient: B = 0 when t_m = 0 or when citations grow linearly with time, and B is non-positive for concave citation trajectories.The index therefore distinguishes delayed growth from trajectories that remain on or above the reference line.
  • Beauty coefficient: B increases with sleeping-period length and awakening intensity, uses the entire citation history, and penalizes early citations.At equal total citations, later accumulation produces a higher B value.
  • Awakening time: The awakening time t_a is the paper age maximizing the distance between (t, c_t) and the reference line ℓ_t.The figure illustrates this distance-based definition together with the beauty coefficient.

B. Awakening time

The awakening time identifies when a paper’s citation trajectory departs most from its reference line, providing a way to locate abrupt citation changes. Applied to APS and WoS data, the measure captures both classic and less extreme delayed-recognition patterns.

  • Awakening time: The awakening time t_a is the age at which the distance between a citation point and the reference line is maximal.The definition is intended to capture the year when an abrupt change in citation accumulation occurs.
  • Awakening time: The distance-based definition works for limit cases with no citations until a later spike and captures the qualitative notion of awakening time.Around awakening, SB co-citation dynamics also exhibit clear topical patterns.
  • Datasets: The analysis uses 384,649 APS papers and 22,379,244 WoS papers with at least one citation, spanning more than a century.APS represents a monodisciplinary physics dataset, whereas WoS includes natural and social sciences for multidisciplinary analysis.
  • Validation: Among 12 revived classics identified by Redner, 6 appear in the top 10 according to B, while the other 6 have very high B values.The agreement supports B’s ability to identify established examples of delayed recognition.
  • Sleeping Beauties in physics: The highest-B APS papers show long hibernation followed by citation bursts without requiring extremely high citation counts.B therefore ranks delayed-recognition trajectories beyond only the most-cited papers.
  • Sleeping Beauties in physics: Top APS Sleeping Beauties form coarse topical groups whose papers have remarkably similar citation histories.The results also include the EPR paradox paper among the highest-B WoS examples.

B. How rare are Sleeping Beauties?

Beauty coefficients form heterogeneous but continuous distributions rather than separating papers into ordinary and exceptional classes. This pattern supports treating delayed recognition as a spectrum, while comparisons across paper ages require caution.

  • Distribution of beauty coefficients: B does not rely on arbitrary thresholds for sleeping time, citation count, or age, enabling analysis of the phenomenon at the systemic level.The measure is designed to quantify delayed recognition without predefined cutoff values.
  • Distribution of beauty coefficients: APS and WoS show heterogeneous but continuous beauty-coefficient distributions, with similar shapes apart from WoS’s larger cutoff.The distributions span several orders of magnitude.
  • Distribution of beauty coefficients: Most papers have low B, but a consistent number have high B, and the distributions show no typical value or mode.There is no clear demarcation separating Sleeping Beauties from normal papers.
  • Distribution of beauty coefficients: Delayed recognition occurs across a wide continuous range, contrasting with claims that Sleeping Beauties are extraordinary cases.The most extreme cases are presented as part of broader heterogeneous distributions.
  • Caveat: Beauty coefficients are not fully straightforward to compare across paper ages because later papers have had less opportunity to develop long sleeping periods.This age-related constraint limits direct comparison between newer and older papers.
  • Distribution of beauty coefficients: In the APS distribution, a power-law fit has exponent α = 2.35 and minimum fitted value B_m = 22.27.These estimates describe the fitted tail using the stated statistical procedure.
  • Citation trajectories: Nearly 90% of papers experience a drastic decline after their maximum yearly citation count, irrespective of B.The empirical distributions remain essentially unchanged when restricting analysis to papers with this typical post-maximum decline.

C. Is the Sleeping Beauty phenomenon statistically significant?

The observed Sleeping Beauty distributions are unlikely to arise from simple citation-network baselines alone. Randomized and preferential-attachment models produce citation histories and beauty-coefficient distributions unlike those of empirical Sleeping Beauties.

  • Model comparison: The paper tests whether observed beauty-coefficient distributions can be explained by idealized network-evolution models, including randomization and preferential attachment.These models serve as baselines for assessing the statistical significance of the Sleeping Beauty phenomenon.
  • Randomization baseline: Randomized citation histories typically decline rapidly, whereas empirical Sleeping Beauties show delayed awakening patterns.The randomization preserves time order while reshuffling citations.
  • Age effects: The probability that older papers receive citations decreases as more later papers become potential recipients, reducing typical beauty coefficients.The citation opportunity structure changes over time as the literature grows.
  • Preferential-attachment baseline: Preferential attachment produces slowly increasing yearly citations through positive feedback but only a small overall citation count.Its beauty-coefficient distribution has a narrower range and a well-defined cutoff.
  • Model comparison: Compatibility with a recently proposed citation-history model remains unresolved.The authors explicitly leave this comparison for further assessment.

D. Sleeping Beauties in science

Extreme Sleeping Beauties occur across physics, chemistry, mathematics, statistics, medicine, social sciences, and multidisciplinary science. The ranking includes papers with very long sleeping periods and high beauty coefficients, showing that delayed recognition is not confined to physics.

  • Top Sleeping Beauties: Statistics and probability emerge as prominent sources of Sleeping Beauties, including a paper that slept for more than one century.The 1901 paper by Karl Pearson relates principal component analysis to minimization chi-distance.
  • Top Sleeping Beauties: Sleeping Beauties include influential techniques such as Fisher’s exact test, the Metropolis–Hastings algorithm, and Kendall rank correlation coefficient.These examples have high beauty coefficients.
  • Disciplinary distribution: Sleeping Beauties are found in the social sciences, contrary to previous claims that they were allegedly absent or rare there.The paper reports numerous social-science examples.
  • Disciplinary distribution: Physics, chemistry, and mathematics are top disciplines producing Sleeping Beauties, while medicine, statistics and probability also appear among the leading categories.The multidisciplinary sciences category ranks third and includes journals such as Nature, Science, and PNAS.
  • Multidisciplinary science: Multidisciplinary journals may attract contributions that become field-defining decades after publication.The authors connect delayed recognition in these venues with contributions perceived as premature or futuristic.

E. What triggers the awakening of an SB?

Awakening can occur when a paper is discovered by another discipline and incorporated into influential work. The examples show citation contexts and terminology changing substantially after the paper becomes relevant to a new research community.

  • Garfield example: The 1955 Garfield paper slept for almost 50 years before becoming suddenly popular around 2000.Its delayed recognition was triggered by later articles by the same author and subsequent influential work.
  • Garfield example: Garfield’s paper was later cited in network science and in debates over the limitations and misuse of journal impact factors.The relevant examples are Kleinberg’s HITS algorithm paper and Seglen’s impact-factor critique.
  • Garfield example: The changing context of Garfield’s paper is visible in citing-paper keywords, with “impact factor” becoming the main post-2000 distinction.Keyword clouds compare citing papers published before and after 2000.
  • Zachary example: Zachary’s 1977 paper remained essentially unnoticed for about 30 years before becoming important in network science.Its social network became a benchmark for validating community-detection methods after Girvan and Newman’s seminal paper.
  • Interdisciplinary awakening: For about 80% of top Sleeping Beauties, at least 75% of citations come from disciplines different from the cited paper’s discipline.Top Sleeping Beauties have a typically much higher cross-disciplinary citation fraction than the comparison categories.

III. DISCUSSION

The paper introduces a parameter-free measure for quantifying delayed recognition and finds that Sleeping Beauties occupy a heterogeneous but continuous spectrum rather than forming exceptional outliers. It also identifies limitations in comparing beauty across disciplines or publication ages and notes that the mechanisms behind awakening remain unresolved.

  • III. DISCUSSION: A parameter-free method quantifies the extent to which a paper can be considered a Sleeping Beauty.The method is applied through a systematic analysis of large-scale bibliographic databases over observation windows longer than a century.
  • III. DISCUSSION: Comparing beauty coefficients across disciplines or paper ages may be problematic because overall citation patterns differ.This limitation affects cross-group comparisons of the measure.
  • III. DISCUSSION: Papers with long dormant periods followed by rapid citation growth are extreme cases within heterogeneous but continuous distributions, not exceptional outliers.The conclusion is supported across the analyzed citation histories.
  • III. DISCUSSION: The empirical distributions of beauty coefficients are not easily reconciled with simple cumulative-advantage models.Those models can reproduce overall citation distributions but not the observed beauty-coefficient distributions.
  • III. DISCUSSION: Further work is needed to uncover the general mechanisms responsible for Sleeping Beauty awakening.The paper identifies the phenomenon statistically but leaves its general mechanisms for future research.

Supporting Information

The supporting analysis uses APS and Web of Science bibliographic datasets spanning more than a century. It restricts analysis to papers with at least one citation and reports the resulting dataset sizes and observation-period effects.

  • Supporting Information: The APS dataset contains 463,348 papers published from 1893 to 2009 in APS journals.The dataset is publicly available upon request.
  • Supporting Information: The Web of Science dataset contains 35,174,034 papers published between 1900 and 2011 across most research fields.The database is available upon purchase from Thomson Reuters.
  • Supporting Information: The APS dataset records only citations originating from papers within APS journals, so it contains fewer citations than Web of Science.Most APS papers are also included in Web of Science.
  • Supporting Information: Among papers receiving at least one citation, the analysis includes 384,649 APS papers and 22,379,244 Web of Science papers.These counts define the cited-paper samples used in the analysis.
  • Supporting Information: The yearly number of cited papers decreases sharply near the observation-period endpoint because recent papers have had less time to accumulate citations.The supporting figure tracks papers with at least one citation received before the observation period ended.

S2. EXAMPLES OF TOP SLEEPING BEAUTIES

The supporting information documents examples of top Sleeping Beauties in APS and Web of Science and provides discipline-specific tables and citation histories for leading cases.

  • S2. EXAMPLES OF TOP SLEEPING BEAUTIES: The citation histories of the top 24 Sleeping Beauties in the APS dataset are shown in supporting figures.Their basic comparison with Redner’s results is provided in a supporting table.
  • S2. EXAMPLES OF TOP SLEEPING BEAUTIES: The citation histories of the top 15 Sleeping Beauties in the Web of Science dataset are displayed in a supporting figure.These cases correspond to the top Sleeping Beauties reported in the main text.
  • S2. EXAMPLES OF TOP SLEEPING BEAUTIES: Supporting tables present basic information on top Sleeping Beauties in Statistics, Mathematics, and Social Sciences and Humanities.Corresponding citation histories are provided in additional supporting figures.

S3. CHARACTERIZING DECREASING PATTERNS

This section characterizes how yearly citations decline after their peak and classifies papers according to whether citations fall below half their maximum during the observation period. Most papers decline quickly after peaking, while many recent awakenings remain unresolved at the endpoint.

  • S3. CHARACTERIZING DECREASING PATTERNS: For most papers, the yearly citation rate decreases quickly, possibly exponentially, after its peak.The observed rapid decline is summarized across the analyzed papers.
  • S3. CHARACTERIZING DECREASING PATTERNS: The analysis considers 189,673 APS papers and 14,689,643 Web of Science papers with positive beauty coefficients.These represent 49.3% of cited APS papers and 65.6% of cited Web of Science papers, respectively.
  • S3. CHARACTERIZING DECREASING PATTERNS: About 60% of papers receive their maximum yearly citations in the last year of the observation period.These are treated as recently awakening papers whose later decline cannot yet be observed.
  • S3. CHARACTERIZING DECREASING PATTERNS: A paper’s half-life is the number of years required for yearly citations to decrease from their maximum to half that maximum.The measure is defined only for papers whose yearly citations fall below ctm/2 after the peak.
  • S3. CHARACTERIZING DECREASING PATTERNS: Yearly citations of Sleeping Beauties decrease rapidly after the peak regardless of their beauty-coefficient ranking.The pattern is reported for the top 1%, the 1%-to-10% group, and the remaining papers, and is confirmed in Web of Science.

S4. NULL MODELS

The analysis compares citation-network randomization and preferential attachment as null models for testing whether beauty coefficients reflect citation dynamics beyond network structure. Randomization preserves degree information while destroying yearly citation dynamics, whereas preferential attachment generates citations through cumulative citation-based attachment.

  • Two null models are used on the APS dataset: citation network randomization and preferential attachment.
  • Citation network randomization: Network randomization swaps citation endpoints while preserving each paper’s reference count and total citation count.Swaps are constrained to avoid shared source or target nodes, duplicate links, and citations violating publication-year ordering.
  • Citation network randomization: The randomized network retains in- and out-degree but destroys the dynamics of yearly citations.
  • Preferential attachment: Preferential attachment adds papers over time with empirical publication counts and assigns references with probability proportional to one plus prior citations.The model begins from the empirical APS network and uses each paper’s reference count when adding new papers.

S5. COARSE TOPICS OF SLEEPING BEAUTIES IN THE APS

High-beauty papers in the APS dataset form coarse topical groups whose members often share citation histories and awakening patterns. The citation network also connects delayed-recognition papers across topics, including double exchange, quantum mechanics, graphite, and graphene.

  • The citation network of the 100 APS papers with highest B values reveals coarse topics and weakly connected components.Node size represents total citations, while isolated nodes are omitted.
  • Papers within the same group often exhibit remarkably similar citation histories, awakening in the same year with similar rising and declining patterns.
  • One identified group concerns the double exchange mechanism, while another concerns quantum mechanics and centers on the EPR paradox paper.
  • The quantum-mechanics group is linked to a theory introduced in the 1950s that became popular in the 1990s.
  • The analysis also identifies a group with complex fluctuations in citation histories.
  • A graphite-and-graphene group centers on pioneering work on graphite’s band structure, foundational to graphene’s discovery.The passage links graphene’s discovery to the subject of the 2010 Nobel Prize in Physics.
  • Supplementary analyses report top Sleeping Beauties across physics, Statistics, Mathematics, and Social Sciences and Humanities using B-based rankings.The supplementary figures and tables provide citation histories, awakening years, B values, and ranked papers for these datasets and fields.
  • For citation-history analyses, awakening years are marked by vertical red lines and yearly citations are compared with network-randomization and preferential-attachment models.The APS analysis ends in 2009, while the WoS analyses cited here end in 2011.
Loading 1505.06454v1…