Source-linked AI summary

A Review of Theory and Practice in Scientometrics

John Mingers, Loet Leydesdorff

arXiv:1501.05462v3cs.DL

TL;DR

Scientometrics addresses how to quantify and evaluate science as a communication system, especially through citations, amid uneven coverage and concerns about what metrics represent. This review synthesizes the field’s history, data sources, indicators, normalization, journal metrics, mapping, evaluation, policy, and future developments. It highlights skewed citation distributions, substantial variation across fields and journals, and limitations showing that citations are not synonymous with quality.

  • Problem

    Scientometrics needs measures for evaluating scientific research, but citation coverage varies substantially across fields and citations do not necessarily represent quality.

  • Method

    The review synthesizes scientometrics’ historical development, citation data sources, metrics, normalization, journal indicators, mapping, evaluation, policy, and future directions.

  • Results

    Citation distributions are highly skewed, with zero-citation proportions ranging from 5% in Management Science to 22% in Omega.

  • Takeaways & Limitations

    Scientometric evaluation requires attention to field differences, database coverage, and the distinction between citation counts and research quality.

  • Takeaways & Limitations

    Citation coverage is very good in the natural sciences, moderate and variable in the social sciences, and generally poor in the arts and humanities.

Abstract

from arXiv · show

Scientometrics is the study of the quantitative aspects of the process of science as a communication system. It is centrally, but not only, concerned with the analysis of citations in the academic literature. In recent years it has come to play a major role in the measurement and evaluation of research performance. In this review we consider: the historical development of scientometrics, sources of citation data, citation metrics and the "laws" of scientometrics, normalisation, journal impact factors and other journal metrics, visualising and mapping science, evaluation and policy, and future developments.

1. HISTORY AND DEVELOPMENT OF SCIENTOMETRICS

Scientometrics developed as the quantitative study of science as an informational and communicative process, centered in practice on citations and their links among research outputs. Its history expanded from citation indexing and bibliometric measures toward web-based metrics, broader databases, research evaluation, and awareness that metrics can shape the realities they measure.

  • Scientometrics studies quantitative aspects of science as an informational process, while bibliometrics applies mathematical and statistical methods to books and other media.
  • Webometrics and altmetrics broadened measurement beyond conventional citations to web resources and online activity such as views, downloads, likes, blogs, and Twitter.
  • Citations became the field’s core empirical resource because they link people, ideas, journals, institutions, and publications across time.
  • Garfield’s Science Citation Index began as a literature-search tool and later supported empirical studies of science, field mapping, research evaluation, and policy.
  • The field developed alongside citation databases, journals, conferences, research units, and policy applications, including the impact factor as a longstanding journal-evaluation measure.
  • Expanded database coverage improved access but left social sciences, humanities, and books unevenly represented, while different databases could produce different results.
  • Scientometric methods increasingly influence the academic realities they measure rather than merely reflecting or mapping a pre-existing reality.

2. SOURCES OF CITATIONS DATA

Citation data come primarily from Web of Science (WoS), Scopus, and Google Scholar (GS), whose coverage, accessibility, citation counts, and data quality differ substantially. These differences make database choice especially consequential across disciplines and for individual metrics.

  • Major sources: WoS and Scopus are traditional subscription-based citation databases, while GS is a later, free alternative that also covers books and conference proceedings.Scopus covers about 20,000 journals; WoS covers around 12,000 journals, conference proceedings, and increasingly books.
  • Database comparisons: GS generally covers about 90% of research outputs across subjects and produces two to five times as many citations for particular works as competing databases.Its broader source range includes materials beyond the journals indexed by WoS and Scopus.
  • Database comparisons: GS coverage is broader but its data quality is poorer because duplicate entries, spelling or date differences, and non-research sources generate incorrect or duplicated records.One psychology comparison found 16.5% incorrect citations in GS, versus 1% or less in the other sources.
  • Metric sensitivity: The same publication set can yield an h-index from 21 to 48 and cites per paper from 10.8 to 56.2, depending on the data source and inclusion rules.Alternative searches identified 88, 349, or 316 papers and substantially different citation totals before calculating these ranges.

3. METRICS AND THE “LAWS” OF SCIENTOMETRICS

This section reviews productivity and citation-impact indicators alongside recurring statistical laws used to describe scientific output and citation patterns. It also examines skewed citation distributions, temporal citation dynamics, and limitations of citation-based evaluation.

  • Scientometric analysis covers indicators of research productivity and citation impact.
  • Productivity laws: Lotka’s Law describes author productivity, with authors making n contributions numbering about 1/n^2 of those making one contribution.
  • Journal and word distributions: Bradford’s Law divides journals into zones with similar article counts, while the number of journals required grows with a power law.
  • Journal and word distributions: Zipf’s distribution relates word frequency inversely to rank, and related productivity, journal, and keyword patterns can arise from cumulative advantage.
  • Limits and extensions: Lotka’s distribution alone is too simplistic because productivity varies over time and subject; a negative-binomial mixture provides a good empirical fit and yields SBS as a consequence.The mixture models publication counts with a Poisson process whose parameter varies according to factors such as age, activity, and discipline.
  • Citation-impact indicators: Citation counts are highly skewed: among six OR journals, papers with zero citations ranged from 5% in Management Science to 22% in Omega.The observed distributions also had modes at zero in all journals except Management Science.

4. Normalisation Methods

Citation rates differ substantially across disciplines and subfields, so meaningful evaluation requires normalization against appropriate field, time, or source-reference contexts. The review compares field-classification, source-normalization, and percentile-based approaches, noting both methodological advantages and statistical limitations.

  • Motivation: Citation rates vary sharply across fields, making cross-field comparisons of researchers, journals, or institutions meaningless without normalization.Mean citation rates were ten times higher in molecular biology than computer science, while management and strategy papers averaged nearly four times as many citations as public administration.
  • Field Classification Normalisation: Field-classification normalization compares received citations with worldwide expectations for the relevant field and publication date.The Leiden methodology compares a unit’s citations with expected citations across the appropriate field and time period.
  • Field Classification Normalisation: The traditional crown indicator has been criticized because its calculation order and mean-based aggregation can misrepresent highly skewed citation distributions.The alternative procedure divides citations by expected values for each paper before averaging the resulting ratios.
  • Field Classification Normalisation: The newer mean normalised citation score (MNCS) is theoretically preferable to the older method, although empirical comparisons found little practical difference.The older method can overweight highly cited fields and produce inconsistent institutional rankings when publication and citation improvements are equal.
  • Source Normalisation: Source normalization defines reference sets from papers citing the evaluated collection, avoiding ad hoc field categories and accommodating interdisciplinary journals.Audience factor, fractional counting, and revised SNIP weight citations according to the reference practices of citing journals or papers.
  • Source Normalisation: Comparative evidence favored source-normalization methods over field classification, especially for interdisciplinary journals, with audience factor and revised SNIP performing best.Fractional counting did not fully match the stronger performance of the leading source-normalization methods.
  • Percentile-Based Approaches: Percentile-based approaches summarize skewed citation distributions by combining the proportions of papers at different citation levels into a single value.The review notes that a single summary remains useful despite the underlying distribution of papers across performance levels.

5. Indicators of Journal Quality: The Impact Factor and Other Metrics

Journal metrics extend beyond the widely used JIF to account for citation longevity, journal prestige, total influence, and field differences. The review emphasizes that JIF is easy to understand but limited by field dependence, short windows, opaque calculation, skewed distributions, and possible manipulation.

  • Journal Impact Factor: Cited half-life measures the median age of papers cited by a journal, indicating how long its citations persist.A five-year cited half-life means that half of citations refer to papers published within the previous five years.
  • Journal Impact Factor: The JIF is the mean citations per paper for a journal over a two- or five-year window.A 2014 two-year JIF counts citations in 2014 to papers published in 2012 and 2013, divided by the number of those papers.
  • JIF Limitations: JIF values differ greatly across fields, so comparing them across disciplines is inappropriate without normalization.The top cell-biology journal had a JIF of 36.5, compared with 7.8 for the top management journal; the twentieth values were 9.8 and 2.9.
  • JIF Limitations: The JIF’s two-year window is too short for many disciplines, particularly where cited half-lives are long.Management journals often have cited half-lives exceeding 10 years, whereas cell biology is typically below 6 years; the five-year JIF is better in this respect.
  • JIF Limitations: JIF calculation lacks transparency, depends on denominator choices and database coverage, and can produce differing values across sources.Studies reported difficulties reproducing figures and differences between WoS and Scopus JIFs for economics because journal coverage differs.
  • JIF Limitations: Because citation distributions within journals are highly skewed, JIF does not represent individual papers or researchers well.A few highly cited papers can inflate the mean while many papers may never be cited.
  • Other Metrics: Prestige-weighted metrics use iterative citation-network algorithms, while Eigenfactor and article influence score capture journal influence in different ways.Eigenfactor follows citation paths through journals and is affected by the total number of papers, whereas article influence normalizes Eigenfactor by journal paper share.
  • Other Metrics: A study found that journals with high JIFs but lower SJRs gained citations from less prestigious sources, showing that citation counts and citation prestige can diverge.The review reports that JIF and SJR were highly correlated overall, but SJR exposed this difference in citation sources.

6. Visualizing and mapping science

Scientometric visualization represents science as networks of relations among words, journals, citations, and other scholarly entities. The review contrasts mapping methods and illustrates how visualizations expose a field’s current topics, knowledge base, and disciplinary environments.

  • Development of mapping approaches: Citation, co-citation, and co-word relations support dynamic network views of scientific development.Co-citation and co-word analysis emerged in the 1970s, while journal-citation data enabled systematic visualization from the mid-1980s.
  • Methods for constructing maps: Multidimensional scaling maps similarity data into lower-dimensional space, whereas force-based methods model network topology through weighted links and energy optimization.MDS uses distance measures such as Euclidean distance or correlation-based similarity; spring-embedded approaches emphasize graph relationships and geodesic distances.
  • Methods for constructing maps: Cosine-normalized matrices can be treated as vector-space distances, while modularity algorithms identify latent community structures in observable networks.The distance can be defined as 1 − cosine, and modularity Q is normalized between zero and one.
  • Applications: A map of 58 frequently occurring EJOR title words revealed groupings around transportation, optimization, decision analysis, performance measurement, and management.The words occurred at least ten times among 505 EJOR documents published in 2013.
  • Applications: The 613 journals cited by 505 EJOR papers were overlaid on a global science map, showing links to environmental sciences, chemical engineering, and biomedicine.These cited sources operationalize the knowledge base on which the articles draw; Rao-Stirling diversity was low at 0.1187, indicating that specialty-internal citation prevailed.
  • Applications: A local map of 29 highly cited Operations Research journals produced central OR, mathematical, and economics-and-finance groupings.The visualization represents the field’s current state, knowledge base, and relevant environments through successive map types.

7. Evaluation and Policy

Bibliometrics has become important for research evaluation because it can complement costly and opaque peer review with broader quantitative information. Its use requires robust data, suitable field-sensitive metrics, attention to interdisciplinarity and behavioral effects, and transparent combination with peer review.

  • Rationale for bibliometric evaluation: Peer review is time consuming, costly, biased, opaque, and limited in the detailed information it provides.These drawbacks motivate interest in bibliometric approaches to research evaluation.
  • Rationale for bibliometric evaluation: Bibliometrics has the potential to provide a cheaper, more objective, and more informative mode of analysis, but should be combined with transparent peer review.In natural and formal sciences in Italy, bibliometrics was judged superior across accuracy, robustness, validity, functionality, time, and cost.
  • Requirements and limitations: Effective evaluation requires robust comprehensive data, yet major databases have limited coverage in the humanities and social sciences, while Google Scholar is less reliable and transparent.Full bibliometric evaluation is feasible in science and some social-science areas but not broadly in the humanities or some technology fields.
  • Requirements and limitations: Metrics should account for disciplinary differences in citation practices rather than rely solely on simple counts, h-indexes, or journal impact factors.More sophisticated alternatives include source-normalized, fractional, prestige-sensitive, and percentile-based measures, although greater sophistication can reduce transparency and replicability.
  • Interdisciplinarity: Research evaluation may systematically disadvantage interdisciplinary work, while reliable and feasible methods for measuring interdisciplinarity remain under development.The review describes bias against interdisciplinary research in business and management and surveys typologies, citation metrics, and entropy measures addressing it.
  • Behavioral and ethical effects: Measurement changes researchers’ behavior, including increased publication, movement toward highly cited journals, greater collaboration, and pursuit of US-based high-impact journals.The review also warns that emphasizing four-star journals can reduce innovation and reinforce the status quo.
  • Behavioral and ethical effects: Inappropriate use of simple indicators such as the h-index or JIF can ignore their limitations, biases, and ethical implications.The problem may arise from how data and metrics are used rather than from the measures themselves.

8. Future Developments

Future scientometrics is expanding beyond citations through altmetrics and broader theories of scholarly communication. The review argues that quantitative citation analysis and sociological accounts of citing behavior should be connected rather than treated as separate approaches.

  • Altmetrics: Altmetrics broaden impact assessment by tracking activities such as viewing, saving, discussion, recommendation, and citation.These channels include repositories, social media, blogs, F1000Prime, Wikipedia, CrossRef, Web of Science, Scopus, and Google Scholar.
  • Altmetrics: Altmetrics can illuminate scholarly impacts on the general public, although viewing showed only a weak correlation with citations.The review presents this public-facing reach as distinct from impact within the academic community.
  • Altmetrics: Altmetrics remain immature because they can be gamed, lack explanatory theory, may reward controversy or fashion, and do not yet cover all impact channels.Most papers also have little social-networking presence.
  • Theories of citation: Network approaches treat citations, words, and scholarly texts as relations that can be visualized or formalized through matrices and latent dimensions.Within a communication-system view, a paper and its publication provide information and utterance, while future references provide understanding.
  • Theories of citation: Scientometrics contains a bifurcation between quantitative analysis of citation events and sociological theorizing about citing behavior.The review states that these perspectives need to be linked because behavior generates citation events, while citation patterns can inform theories of scientific communication.
  • Theories of citation: A communication-system perspective models science as recursively connected communications operating at an emergent level distinct from individual scientists.The review frames scientometric visualization as a way to reveal cognitive distinctions generated by these systems.
Loading 1501.05462v3…