Source-linked AI summary

OpenCitations, an infrastructure organization for open scholarship

Silvio Peroni, David Shotton

arXiv:1906.11964v3cs.DL

TL;DR

Scholarly citation data are often inaccessible or restricted for reuse, limiting equitable access and reproducible bibliometric research. OpenCitations responds with open Linked Data infrastructure, datasets, identifiers, and services built around Semantic Web technologies. It has released more than 450 million citations, while its future development expands coverage and metadata; long-term operation depends on continuing support from the scholarly community or adopting institutions.

  • Problem

    Proprietary access and reuse restrictions limit equitable access to citation data and the reproducibility of research studies using them.

  • Method

    OpenCitations provides global citation data and related services as open Linked Data, using Semantic Web technologies, persistent identifiers, and citation-centered indexes.

  • Results

    More than 450 million citations have been released, with citation data and services receiving more than 3.1 million website accesses from over 68,000 unique visitors in the past year.

  • Takeaways & Limitations

    OpenCitations is expanding open citation coverage through new indexes and OpenCitations Meta, which will add publication metadata currently lacking from the OpenCitations Corpus.

  • Takeaways & Limitations

    Long-term sustainability requires ongoing support because OpenCitations’ free and open model provides no products to sell, membership fees to charge, or freemium income stream.

Abstract

from arXiv · show

OpenCitations is an infrastructure organization for open scholarship dedicated to the publication of open citation data as Linked Open Data using Semantic Web technologies, thereby providing a disruptive alternative to traditional proprietary citation indexes. Open citation data are valuable for bibliometric analysis, increasing the reproducibility of large-scale analyses by enabling publication of the source data. Following brief introductions to the development and benefits of open scholarship and to Semantic Web technologies, this paper describes OpenCitations and its datasets, tools, services and activities. These include the OpenCitations Data Model; the SPAR (Semantic Publishing and Referencing) Ontologies; OpenCitations' open software of generic applicability for searching, browsing and providing REST APIs over RDF triplestores; Open Citation Identifiers (OCIs) and the OpenCitations OCI Resolution Service; the OpenCitations Corpus (OCC), a database of open downloadable bibliographic and citation data made available in RDF under a Creative Commons public domain dedication; and the OpenCitations Indexes of open citation data, of which the first and largest is COCI, the OpenCitations Index of Crossref Open DOI-to-DOI Citations, which currently contains over 445 million bibliographic citations and is receiving considerable usage by the scholarly community.

1. Introduction

Citation data underpin scholarship and reproducible bibliometric analysis, yet major citation indexes remain proprietary or restrict reuse. OpenCitations addresses this gap by providing global scholarly citation data as a fully free Linked Data alternative, alongside related open resources.

  • Why open citations matter: Bibliographic citations connect scholarly works and support credit assignment, making their open availability crucial for reproducible bibliometric and scientometric research.Open citation data also support appraisal of research, reduction of misconduct, and equitable participation in science.
  • Limits of existing indexes: Web of Science and Scopus are authoritative citation sources, but their subscription costs exclude institutions and independent scholars that cannot afford access.Other general citation sources may be free to access while still restricting reuse and republication.
  • OpenCitations: OpenCitations provides a fully free and open alternative for accessing global scholarly citation data in Linked Data form.The article introduces its main data, services, known uses in scientometrics, and planned developments.
  • Related open resources: The NIH Open Citation Collection provides another open citation source, but its coverage is limited to citations between articles indexed in PubMed.Its citations can be retrieved through the iCite web service.

2. From the origins to the Initiative for Open Citations

The OpenCitations Corpus began as an early open bibliographic and citation dataset, but its initial scope was limited. Open access to deposited reference lists later supported broader open-citation efforts.

  • Origins: The OpenCitations Corpus was first released in 2010 as the main output of the JISC-funded Open Citations Project.It was the first dataset released containing open bibliographic and citation data.
  • Early limitations: The initial OpenCitations Corpus had limited scope, despite subsequent reference-list contributions between 2010 and 2016.The supplied passage introduces this limitation while describing the early development period.
  • Toward open citations: OASPA required members depositing reference lists with Crossref to make those lists openly available in line with Initiative for Open Citations recommendations.This policy connected publisher reference-list deposits with broader open-citation availability.

3. The benefits of open scholarship and open citations

Open scholarship expands access to publicly funded knowledge through open information sources, while open citations provide specific benefits for researchers, bibliometricians, and librarians.

  • Changing scholarly communication: The open scholarship model is associated with disruptive changes in scholarly publishing, including the crumbling of subscription models for journal content.The passage cites the growing number of academic libraries, institutions, and consortia cancelling Big Deals.
  • Open scholarship: The open scholarship movement is based on making the fruits of publicly funded scholarly work openly available to scholars and the public.It complements broader open publication of knowledge, including resources such as Wikipedia.
  • Open scholarship: Open scholarship includes open-source software, open-access scientific publication, and open research datasets.The passage gives Linux, PLoS, eLife, F1000 Research, and the Protein Data Bank as examples.
  • Benefits of open citations: Open citations particularly benefit researchers without subscription access, bibliometricians publishing the data behind their findings, and librarians supporting stakeholder communities.The stakeholder groups include authors, researchers, students, and institutional administrators.

4. A brief introduction to Semantic Web technologies

The growth of open scholarship creates a need for scalable, distributed scholarly infrastructure rather than a single centralized database. Semantic Web technologies address this need by making data precise, machine-processable, and interconnected.

  • Infrastructure challenge: A centralized database is not feasible long term for growing bibliographic and citation data because of performance and physical-space demands.The required infrastructure should support access, REST APIs, and querying appropriately as data volumes increase.
  • Infrastructure challenge: COAR calls for a distributed globally networked scholarly infrastructure with cross-repository connections, bidirectional links, and machine-friendly data formats.These features are intended to support a wider range of global repository services.
  • Semantic Web technologies: W3C Semantic Web standards facilitate precise semantics for encoding data and making it machine-processable on the Web.They provide a technological basis for the distributed scholarly infrastructure described in the section.
  • Semantic Web technologies: Semantic Web data identify entities, types, and relationships with HTTP URIs, use shared ontologies, and represent statements as RDF subject-predicate-object triples.The example <Paper A> cito:cites <Paper B> illustrates a citation relationship encoded as an RDF triple.

5. OpenCitations: tools and services

OpenCitations provides open infrastructure for modeling, identifying, publishing, and accessing bibliographic and citation data. Its data model, persistent identifiers, datasets, and indexes support machine-readable description and analysis of citations.

  • OpenCitations publishes open bibliographic and citation data using Semantic Web technologies and advocates for open citations.
  • Open Citation Identifiers: Open Citation Identifiers are globally unique persistent identifiers that encode citing and cited publications and their supplier databases.The OCI resolution service retrieves metadata in RDF, Scholix, JSON, or CSV formats.
  • Open Citation Identifiers: Treating citations as first-class data entities centralizes their metadata and makes them easier to describe, distinguish, count, process, and analyze.Aggregated citation data can support citation-network visualization and analysis of citation time spans by discipline.
  • The OpenCitations Data Model and the SPAR Ontologies: The OpenCitations Data Model represents bibliographic and citation entities, attributes, and relations in RDF through the SPAR Ontologies.The model covers published resources, manifestations, bibliographic references, and responsible agents.
  • The OpenCitations datasets: The OpenCitations Corpus provides downloadable RDF data under CC0, while its indexes publish open citation links from third-party databases.The OCC contains about 14 million citation links to over 7.5 million cited resources, and COCI contains more than 445 million citations under CC0.
  • The OpenCitations datasets: CROCI enables third parties to submit citation data that are not otherwise openly available, while the indexes describe citation properties and retrieve entity metadata through APIs.Index records can include creation date, citation timespan, and citation type, while bibliographic metadata are retrieved on demand.

6. The sustainability of OpenCitations

OpenCitations pursues long-term sustainability through open infrastructure principles, reusable software and data, broad community engagement, and diversified funding. Because its services are free, continued operation depends on external support and institutional adoption.

  • 6.1. Insurance: OpenCitations follows open infrastructure principles covering insurance, governance, and sustainability.These principles are intended to preserve the work if OpenCitations itself ceases to exist.
  • 6.1. Insurance: Its software, models, and data use open licenses and standards that enable reuse, migration, and adaptation by third parties.Software is released under the ISC License, models under CC-BY, and data under CC0.
  • 6.2. Governance: OpenCitations develops generic applications and citation data covering the full scholarly research domain for reuse in diverse scenarios.Applications include OSCAR, LUCINDA, and RAMOSE.
  • 6.2. Governance: OpenCitations is directed by the article’s authors and administratively managed by the University of Bologna’s independent Research Centre for Open Scholarly Metadata.The Research Centre has an international board drawn from bibliographic stakeholders.
  • 6.3. Sustainability: OpenCitations has relied on project grants and plans to pursue targeted funding from additional foundations and organizations.Its funding history includes JISC and the Alfred P. Sloan Foundation, with later support from Wellcome and planned H2020 participation.
  • 6.3. Sustainability: Because its data and services are free, OpenCitations cannot rely on sales or membership fees and instead requires continuing community or institutional support.Proposed models include SCOSS-style crowd funding, adoption by scholarly libraries, philanthropy, or academic funding agencies.

7. Usage statistics

OpenCitations recorded substantial use across its website, services, countries, and downloadable resources during the past year. Usage measurements distinguish access channels, exclude automated agents, and show international demand.

  • Website and services: More than 3.1 million website accesses came from over 68,000 unique visitors in the past year.Automated agents and bots were excluded from the counts.
  • Website and services: Figure 2 separates access into HTTP content negotiation, interfaces, REST APIs, SPARQL, and other website visits.The monthly counts cover April 2018 through March 2019, and the y-axis is logarithmic.
  • Geographic distribution: Italy, Poland, and the United States generated the most requests, followed by Brazil, France, the Netherlands, Spain, China, Germany, India, and the UK.Figure 3 organizes requests by country using request IP addresses.
  • Figshare resources: More than 20,000 views and 3,000 downloads were recorded for OpenCitations resources on Figshare.These resources include dataset dumps and definition documents.

8. Adoption of OpenCitations by the community

Researchers and software tools have adopted OpenCitations data for bibliometric studies, citation-network visualization, and collaborations involving bibliographic and citation-data projects.

  • Research use: OpenCitations data support open bibliometric research and enable publication of the underlying data for reproducibility.The data are particularly valuable to researchers without subscription access to Web of Science or Scopus.
  • Research use: COCI data have supported studies of PLOS ONE citations, Italian Scientific Habilitation outcomes, and the role of books in scholarly communication.These examples demonstrate use of open citation data in distinct bibliometric investigations.
  • Software adoption: VOSviewer, Citation Gecko, OCI Graphe, and VisualBib use OpenCitations data to construct or visualize citation networks.VOSviewer accesses COCI through its REST API, while other tools retrieve COCI or OCC data for network generation.
  • Academic collaborations: OpenCitations collaborates with academic projects to promote its data model and provide publication venues for liberated citation data.The collaborations include the Venice Scholar Index and other bibliographic-data initiatives.

9. Conclusions and future developments

OpenCitations aims to make scholarly bibliographic and citation data freely usable, while expanding datasets and organizing them into interoperable repositories. Future work includes new citation indexes, richer metadata, and federated infrastructure.

  • Conclusions: OpenCitations has released more than 450 million citations and plans to expand existing data and create new datasets.Its stated goal is to let anyone use open scholarly data and related services for any purpose.
  • Future indexes: New indexes will cover open Wikidata, DataCite, and Dryad citations, substantially broadening the citation-data coverage.These indexes are named WOCI, DOCI, and DROCI.
  • Future datasets: OpenCitations Meta will add abstracts, keywords, author affiliations, and funding details that are currently missing from the OpenCitations Corpus.The database is intended to support richer bibliometric analyses and more complex API calls.
  • Future architecture: The planned New Corpus will use federated SPARQL-based repositories encoded with the OpenCitations Data Model.Separate repositories can describe different data types while remaining interoperable through shared standards.
  • Future architecture: The original OpenCitations Corpus will remain an experimental sandbox for testing software and data-model extensions.It will retain data types together over a finite collection of papers and references.
  • Future architecture: Interoperable repositories could support broader federation with similar third-party open resources.Potential integrations include repositories describing publication types and related scholarly metadata.

Competing Interest Statement

The authors disclose that they are the Directors of OpenCitations, the subject of this paper.

  • The authors are OpenCitations' Directors, and OpenCitations is the subject of the paper.
Loading 1906.11964v3…