Source-linked AI summary
A new methodology for constructing a publication-level classification system of science
Ludo Waltman, Nees Jan van Eck
TL;DR
Journal-level classification systems provide limited detail and struggle with multidisciplinary journals. The paper introduces hierarchical publication-level clustering based on citation relations and applies it to almost ten million publications. The methodology is transparent and relatively simple, but relying exclusively on direct citations limits coverage and accuracy.
Problem
Existing journal-level classification systems offer limited detail and have difficulties with multidisciplinary journals.
Method
The methodology clusters individual publications into hierarchical research areas using citation relations.
Results
Almost ten million publications were included in the application, covering essentially all Web of Science-indexed publications in a ten-year period.
Takeaways & Limitations
The methodology is transparent, relatively simple, and has fairly modest computing and memory requirements.
Takeaways & Limitations
Exclusive reliance on direct citation relations can reduce coverage and accuracy, and other relations could improve both.
Abstract
from arXiv · showhide
Classifying journals or publications into research areas is an essential element of many bibliometric analyses. Classification usually takes place at the level of journals, where the Web of Science subject categories are the most popular classification system. However, journal-level classification systems have two important limitations: They offer only a limited amount of detail, and they have difficulties with multidisciplinary journals. To avoid these limitations, we introduce a new methodology for constructing classification systems at the level of individual publications. In the proposed methodology, publications are clustered into research areas based on citation relations. The methodology is able to deal with very large numbers of publications. We present an application in which a classification system is produced that includes almost ten million publications. Based on an extensive analysis of this classification system, we discuss the strengths and the limitations of the proposed methodology. Important strengths are the transparency and relative simplicity of the methodology and its fairly modest computing and memory requirements. The main limitation of the methodology is its exclusive reliance on direct citation relations between publications. The accuracy of the methodology can probably be increased by also taking into account other types of relations, for instance based on bibliographic coupling.
1. Introduction
The paper proposes a publication-level science classification methodology to address the limited detail and multidisciplinary-journal difficulties of journal-level systems. It clusters publications by citation relations into hierarchical research areas and applies it to almost ten million publications.
- Proposed methodology: The proposed methodology clusters individual publications using citation relations and assigns them to hierarchical research areas.Each publication is assigned to a single research area, from broad disciplines to small subfields.
- Paper scope: The methodology is designed to cluster very large numbers of publications, producing an application covering almost ten million publications.The paper analyzes the methodology’s strengths and limitations after presenting this application.
- Existing classification systems: Journal-level systems such as Web of Science and Scopus assign publications to research areas indirectly through their journals.Web of Science is described as the most popular system and contains about 250 subject categories.
- Motivation: Journal-level classification systems typically offer limited detail and have difficulties handling multidisciplinary journals.
2. Methodology
The methodology determines publication relatedness from direct citations, clusters publications hierarchically, and labels areas from titles and abstracts. Normalization, parameterized resolution, reassignment, and approximate optimization support large-scale classification while imposing stated limitations.
- Overview: The methodology has three steps: determining publication relatedness, clustering publications into research areas, and labeling research areas.
- Step 1: Relatedness: Publication relatedness is based only on direct citations, with citation direction disregarded to reduce memory use and computing time.The resulting binary relation is one when either publication cites the other and zero when no direct citation relation exists.
- Step 1: Relatedness: Relatedness scores are normalized because citation behavior differs substantially among scientific fields.Normalization gives each publication a total normalized relatedness of one with all other publications, ensuring equal overall weight.
- Step 2: Clustering: Publications are clustered bottom-up into hierarchical research areas, with minimum-size rules discarding undersized areas and reassigning their publications.The hierarchy requires lower-level areas to map consistently into higher-level areas.
- Step 2: Clustering: The clustering algorithm approximately optimizes a quality function because exact maximization is usually infeasible for large publication sets.Random elements mean different runs can generally produce different values.
- Step 2: Clustering: Some publications are excluded when their preliminary research areas cannot be properly reassigned because those areas have no relations with eligible areas.
- Step 3: Labeling: Research areas are labeled by extracting terms from member publications’ titles and abstracts and ranking terms by relevance.The relevance calculation balances a term’s relative frequency within a subarea against its absolute frequency there.
3. Application
The application constructs a three-level classification from Web of Science publications published during 2001–2010. Using 10.2 million publications, it produces 20 broad areas, 672 fields, and 22,412 small subfields.
- Dataset: 10.2 million publications from Web of Science were used to construct the classification system for 2001–2010.Included document types were articles, letters, and reviews in the sciences and social sciences; arts and humanities were excluded.
- Results: The resulting system contains 20 research areas at level 1, 672 at level 2, and 22,412 at level 3.
- Computation: The calculations used repeated optimization runs, including 500 runs at the lowest level and 10,000 runs at each of the other two levels.The optimization algorithm was written in C, other calculations in MATLAB, and the calculations used a computer with 64 GB of memory.
4. Results
The publication-level classification system organizes nearly ten million publications into a three-level hierarchy, revealing broad and fine-grained research areas while exposing classification and coverage limitations. Its results include uneven area sizes, identifiable hot fields, and some substantively unsatisfactory assignments linked to reliance on direct citations.
- Coverage: 0.8 million of the 10.2 million starting publications could not be included in the classification system.The excluded publications are overrepresented in earlier years and include a disproportionate share of letters.
- Level 1: 20 level-1 areas range from about 130,000 to almost 1.34 million publications, averaging about 470,000 publications per area.The areas only partially correspond to traditional scientific disciplines, making some labels difficult to determine.
- Level 2: The three hottest level-2 areas are in physics, molecular biology, and virology, including graphene research with 75% of 6,911 publications appearing during 2008–2010.The virology area seems to deal mainly with influenza viruses.
- Level 3: 22,412 level-3 areas average 422 publications, with sizes ranging from 50 to 4,170 publications.The three hottest level-3 areas concern high-temperature superconductivity, graphene, and a crystallography topic; one crystallography publication had more than 20,000 citations despite appearing in January 2008.
- Limitations and validation: Some publications are assigned to areas that are understandable from citation relations but unsatisfactory substantively because the methodology uses only direct citations.The classification can also place JASIST publications outside its two main areas, although some such assignments, including network analysis, appear sensible.
5. Conclusion and future research
The paper presents a publication-level classification methodology that scales to almost ten million publications while remaining transparent and relatively simple. Its main limitation is reliance on direct citations, motivating richer relatedness measures and further work on labeling, overlap, evaluation, and journal-level systems.
- Future research: Automatically generated labels are less satisfactory at higher aggregation levels, where suitable research-area labels are harder to identify.
- Conclusion: The methodology is transparent, relatively simple, fully documented, and based on a limited number of easily understood steps.It uses few manually chosen parameters, and the clustering software is freely available online.
- Conclusion: The methodology has relatively modest computing-time and memory requirements, although large applications may exceed standard desktop capacity.
- Limitations: Exclusive reliance on direct citation relations limits coverage and can assign publications with few citations to incorrect research areas.More sophisticated measures could incorporate indirect citations, bibliographic coupling, or shared words in titles and abstracts.
- Future research: Future work should improve labeling, allow overlapping research areas, evaluate accuracy more rigorously, and explore deriving journal-level systems from publication-level classifications.Rigorous evaluation is difficult because no golden standard is available; expert feedback is suggested.
Appendix
The appendix illustrates how the classification system organizes JASIST publications into research areas at levels 2 and 3. It also lists selected terms and frequently cited publications associated with these areas.
- Level-2 areas: Table A1 lists the five level-2 research areas containing the largest numbers of JASIST publications.The table identifies area 4.30 and its associated terms, including h index, academic library, and document supply.
- Level-3 areas: Table A2 lists the five level-3 research areas containing the largest numbers of JASIST publications.Area 4.30.2 is associated with terms including query, web searching, searcher, and information.
- Frequently cited publications: The appendix provides examples of frequently cited JASIST publications for areas involving link analysis, web links, knowledge organization, and information science.
- Area labels: The appendix also shows that classification-area labels are represented through selected terms such as classification scheme, GIS, and Hjorland.