Source-linked AI summary
Usage Bibliometrics
Michael J. Kurtz, Johan Bollen
TL;DR
Citation analysis has delays, limited coverage, and a narrow view of scholarly activity, creating a need for complementary usage evidence. This review defines and surveys modern usage-based informetrics, examining its collection, analysis, integration with bibliometrics, and scope. It concludes that large-scale usage data can expand bibliometrics substantially, while standardization and aggregation remain unresolved challenges.
Problem
Citation data have publication delays, focus mainly on journal articles and authors, and overlook activities outside the publishing and citation system.
Method
The review examines how usage data are defined, collected, analyzed, combined with traditional bibliometric data, and sampled for representativeness.
Results
Modern usage data can approximate or surpass citation and text databases in scale, quality, detail, community coverage, and request metadata.
Takeaways & Limitations
Usage data offer expanded capabilities for bibliometrics across scholarly activities and phases of the scientific process.
Takeaways & Limitations
Usage data remain difficult to standardize and aggregate because recording formats, fields, semantics, normalization, and privacy-related data loss vary.
Abstract
from arXiv · showhide
Scholarly usage data provides unique opportunities to address the known shortcomings of citation analysis. However, the collection, processing and analysis of usage data remains an area of active research. This article provides a review of the state-of-the-art in usage-based informetric, i.e. the use of usage data to study the scholarly process.
Introduction
Electronic usage data provide detailed records of scholarly activity that complement citation-based bibliometrics, while their interpretation, scope, and privacy implications require care. The review defines and surveys this expanded usage-based approach, focusing on online journal articles.
- Electronic usage records: Electronic libraries record detailed transaction metadata for each resource use, including the user, location, time, request type, record type, and access source.These records are more comprehensive than print-library traces but do not reveal user motivations or goals.
- Why usage data matter: Usage data complement citation and text data because they avoid publication delays, extend beyond journal articles and authors, and cover broader phases of scholarly activity.The authors describe this expanded coverage as powerful but potentially dangerous for bibliometric analysis.
- Expanded bibliometrics: Modern usage data can approximate or surpass citation and text databases in scale, quality, detail, community coverage, and request metadata.They can also support reconstruction of user clickstreams and article-level analysis.
- Study clusters and scope: Usage-based bibliometrics includes user-behavior studies and studies of scholarly actors or units, with privacy and data-sharing concerns affecting user studies.The second cluster can rank articles or authors without identifying individual users.
- Review scope: The review discusses defining, collecting, combining, and sampling usage data, while restricting its scope to online journal articles.Newer scholarly communication forms, including data archives, online databases, workflow systems, and blogs, are left for future work.
Usage Data and Statistics
Usage has been operationalized in multiple ways, including reads, uses, downloads, and hits. A formal definition must accommodate varied contexts without relying on unmeasurable motivations or context-specific details.
- Operational definitions: Usage-based bibliometrics operationalizes usage variously as “reads,” “uses,” “downloads,” or “hits.”The review emphasizes that these terms reflect different operational choices across studies.
- Formal definition: A formal usage definition should cover diverse contexts while excluding unmeasurable user motivations and intentions.It should also avoid context-specific issues that cannot be consistently incorporated.
A Request-Based Model of Usage
The request-based model defines usage as a user’s request for a service concerning a scholarly resource, mediated by an information service. Usage events record these transactions and their associated metadata.
- Model components: The model contains users, information services, and scholarly resources such as books, journal articles, and e-science data sets.The information service mediates requests and returns a service concerning the resource or its representation.
- Definition of usage: Usage occurs when a user requests a service pertaining to a particular scholarly resource from a particular information service.This is the paper’s operational definition of usage.
- Usage events: A usage event is an electronic record of a user-generated request for a resource, service, and point in time.Usage log data are collections of such events over a specified period.
- Interpretive limits: Requests operationalize user interest only indirectly, and request types may express different degrees of interest.For example, abstract viewing is described as indicating less interest than full-text downloading, while the relationship remains under study.
- Recorded metadata: Usage data can include event, user or session, request-type, resource, and date-time identifiers.These fields distinguish events, users or sessions, requested services, resources, and request timing.
Usage Data versus Usage Statistics
Usage data preserve individual request events, whereas usage statistics aggregate those events for a resource and selected parameters. Aggregation supports normalized or unnormalized impact indicators but removes most event-level information.
- Aggregation: An aggregation function can count usage events for a resource within a selected period, producing a usage statistic.The parameters control the aggregation, such as restricting events to a date interval.
- Aggregation model: The formal aggregation maps usage data U to statistic S for resource R under parameter set P.A date range d ∈ P can restrict the statistic to events satisfying d1 < t < d2.
- Data versus statistics: Usage statistics generally omit individual event information but retain the resource identifier, statistic value, and possibly the aggregation period.These retained elements distinguish aggregate statistics from underlying usage data.
- Impact indicators: Usage-derived impact indicators are instances of the aggregation function with parameters producing normalized or unnormalized metrics over a period.Examples include the Usage Impact Factor and Usage Factor, the latter defined as average usage of a journal’s articles over two years.
Usage Data: A Practical Overview
Modern electronic services produce detailed, article-level usage records, but collection remains constrained by incomplete sessions and uneven coverage. Link resolvers and standardized ContextObjects support aggregation across services, while their reach and recorded detail remain limited.
- From Physical to Electronic Usage Data: Electronic usage records capture request details and article-level activity more richly than physical-library statistics.Physical statistics lack request context, individual-article resolution, and broad scale, whereas modern usage data provide more detailed records of individual events.
- Usage Data from Web Server Logs: Web-server logs record HTTP request parameters, but unreliable session information hampers reconstruction of individual request sequences.This limits analysis of clickstreams, resource relationships, and services based on user traffic.
- Link Resolvers: Link resolvers can track requests across multiple OpenURL-enabled scholarly services and provide a common representation framework for usage data.Their hub position lets institutions collect usage across services, while the OpenURL Framework represents request-related fields through a ContextObject.
- Link Resolvers: The OpenURL ContextObject encodes the referent, requester, service type, referring entity, resolver, and referrer associated with a service request.These fields provide a standardized structure for representing usage events and supporting service fulfillment.
- Link Resolvers: Aggregating link-resolver usage data can standardize recording and sharing across communities, but it excludes requests outside participating or OpenURL-enabled services.The approach also omits usage fields not captured by the ContextObject and cannot solve the nonuniversal deployment of linking servers.
Direct Measures
Usage obsolescence reflects multiple user behaviors and varies substantially with the user population, search interface, and document set. These differences complicate direct comparisons between usage and citation measures.
- Usage patterns: Usage obsolescence combines distinct behaviors: browsing new articles, searching current literature, locating specific articles, and approximately age-independent historical use.The model labels these modes N, C, I, and H, respectively.
- Usage patterns: The C, I, and H components fit the 25-year usage data, while the N component is too short-lived to appear in the plotted graphs.The expanded view indicates that all three visible components are necessary.
- User populations: Usage patterns differ sharply across users: professional astronomers show stable modeled obsolescence, whereas Google users show elevated early use followed by a lower, nearly constant level.Google Scholar usage instead rises with article age.
- User populations: Usage obsolescence depends critically on user populations because scholarly-article users are broader and different from scholarly authors.Citation obsolescence is tied to scholarly authors, whereas usage has no equivalent predefined user group.
- Comparison with citations: Comparisons require carefully matched documents because highly used materials can be rarely cited, producing large usage-to-citation differences across journals and article types.The “Astrophysics in XXXX” papers are described as routinely among astronomy’s most used yet very rarely cited.
Some Usage-Based Statistical Measures
Usage data complement citations by capturing current and user-related dimensions of scholarly activity, enabling measures of productivity, usefulness, obsolescence, and country-level research activity. Interpretation requires carefully specified document populations because usage and citation patterns can diverge substantially.
- Usage measures: Usage records contain information about both the entity being used and the user, enabling measures beyond citation or publication counts.Usage-based measures can aggregate article properties or user-based activity, such as downloads from users in a country.
- Caveats: Usage-based evaluation remains limited by differing user and citer populations, possible manipulation of mostly anonymous records, and difficult cross-journal comparisons.Electronic journal design, interfaces, and search methods can change the meaning of usage statistics across journals and publishers.
- Usage measures: Usage rates reflect current use, whereas citation counts reflect integrated past use; together they provide a richer view of scholarly activity.For authors, combining citation counts and usage rates with age yields more information about productivity or usefulness than citation counts alone.
- Countries: Usage data support near-real-time country analyses and show per-capita use related to GDP, while growth patterns differ across countries and regions.The cited examples report quadratic scaling of per-capita astrophysics use with per-capita GDP and continued faster-than-world-average growth for India.
Social Network Measures
Social network measures extend usage bibliometrics beyond simple counts by modeling relationships among journals and weighting the influence or position of their sources. These measures reveal distinctions between popularity, prestige, and interdisciplinary connectivity, while usage- and citation-based rankings can diverge from conventional Impact Factor rankings.
- Scope: Social network analysis models relationships among scholarly entities to assess journal and article impact, and this approach is being extended from citation networks to usage bibliometrics.The review presents network principles, node-status measures, and applications to usage data.
- Degree Centrality: Degree centrality counts incoming or outgoing endorsements, but it indicates popularity rather than necessarily capturing influence or prestige.Because it counts citations without considering their origins, many citations from lower-ranked journals can outweigh fewer citations from prestigious journals.
- Eigenvector Centrality and PageRank: Eigenvector-based measures weight citations from influential journals, addressing the circular problem of defining a journal’s influence through iterative recalculation.PageRank efficiently approximates citation eigenvector centrality for very large, sparse networks.
- Eigenvector Centrality and PageRank: PageRank-based journal rankings can differ substantially from Impact Factor rankings because high citation counts may come from relatively unprestigious sources, while fewer citations may come from prestigious sources.The reported comparison identifies domain-dependent deviations, including review journals in medicine with high Impact Factors but low PageRank scores.
- Shortest Path Measures: Betweenness centrality identifies journals that connect other journals in the network, so removing highly central journals would interrupt paths between many journal pairs.This measure captures an interdisciplinary position in the network structure.
- Comparing Measures: A PCA of 43 journal-impact measures found that the first two components explained nearly 85 percent of variation and separated measures by data source and by popularity versus prestige.Citation- and usage-based PageRank and betweenness measures correlated more strongly with each other than with the Impact Factor.
Open Access
Open-access articles have often shown higher citation rates, but the review finds that this difference is difficult to attribute to open access itself. Evidence instead points to selection and early-access effects, while randomized evidence found increased downloads without increased citations.
- Interpretation: The broader literature has not satisfactorily explained the relationship between open access, usage, and citations.The review notes that correlation does not establish causality and that other mechanisms may account for the observed association.
- Observed citation advantage: Open-access articles have often been cited at roughly twice the rate of comparable restricted articles.This pattern was reported in computer science and across physics subfields covered by arXiv.
- Competing explanations: Historical and comparative analyses found little or no evidence that open access itself caused the citation advantage.Studies instead associated the difference mainly with early access and selection bias.
- Experimental evidence: A randomized trial found significantly more full-text downloads for open-access articles but no difference in citation rates.The design eliminated confusion from early-access and selection-bias factors, although the study was criticized on methodological grounds.
- Beyond citations: Increased usage derived from open access may itself serve as an indicator of increased impact, independently of citation growth.The review also notes easier access for science journalists and bloggers as a potential public benefit.
Mapping of Science from Usage Data
Usage data can extend science mapping beyond citation relations by representing how users move among journals and articles. Early examples show promising possibilities, but representative sampling and longitudinal coverage remain essential requirements.
- Citation-based mapping: Citation-based maps traditionally visualize connections among articles, journals, and scholarly domains.They commonly use co-citation relations and journal citation similarities to identify structure and trends in science.
- Mapping from usage: Usage-based science mapping derives journal relationships from flows of user traffic rather than citation similarities.The approach applies network and principal-component methods to usage graphs and link-resolver records.
- MESUR examples: MESUR researchers extracted journal clickstream maps from 200 million usage events within a 1 billion-event reference dataset.The work was later extended with a larger dataset and validation against subject-classification taxonomies.
- Potential and constraints: Usage-based mapping remains uncommon but could track science as it develops and reveal longitudinal trends relevant to funding agencies and policymakers.Its usefulness depends on meeting minimum requirements for sample characterization and data fidelity.
- Reading the maps: In clickstream visualizations, journals are represented as colored circles and links indicate a high probability that users move from one journal to another.Colors correspond to scientific domains, while links encode sequential usage relationships.
Conclusion
Usage data are becoming central to bibliometrics because online scholarly activity generates detailed records that citation analysis cannot fully capture. Realizing this potential requires standards, representative aggregation, privacy safeguards, and broader acceptance of usage-based measures.
- The changing basis of bibliometrics: Article-level usage records are central to emerging bibliometrics, although their applications have not yet reached the commonplace acceptance of citations.The shift toward online scholarly work is expected to increase their importance.
- Standards and aggregation: Incompatible recording formats and unresolved aggregation practices hinder the standardization of usage bibliometrics.COUNTER and MESUR are cited as projects working toward recording and aggregation standards.
- Privacy and ownership: Usage datasets raise privacy, confidentiality, and ownership concerns for individual users, institutions, and data providers.Anonymous session identifiers can preserve temporal usage patterns while protecting user privacy, but institutional and provider safeguards remain largely ad hoc.
- Coverage beyond citations: Usage data can capture information-rich activity that citation counts miss, including millions of events surrounding dynamic scholarly web services.For systems such as the SkyServer, usage events greatly exceed citations to the papers describing the dataset.
- Sampling and infrastructure: Representative sampling across the scholarly community remains unresolved because no reliable census exists for validating usage samples.Long-term continuity and trusted stewardship are also needed for aggregated indicators.
- Future adoption: The acceptance of usage-based impact measures depends on resolving scientific, logistical, sampling, and interpretive questions.The review specifically asks whether usage measures can become as understandable and simple to apply as established citation measures.
Endnote
The Astrophysical Journal and Astronomy and Astrophysics became fully available online on January 1, 1997; two additional astronomy journals followed on January 1, 1998.
- Online availability: Four astronomy journals became fully available online between January 1, 1997 and January 1, 1998.The Monthly Notices of the Royal Astronomical Society and The Astronomical Journal followed the first two journals on January 1, 1998.