Source-linked AI summary

Which type of citation analysis generates the most accurate taxonomy of scientific and technical knowledge?

Richard Klavans, Kevin W. Boyack

arXiv:1511.05078v2cs.DL

TL;DR

Citation-based taxonomies need objective accuracy evaluation because they inform science-policy decisions, yet literature partitions lack definitive gold standards. This study compares document- and journal-based taxonomies using high-reference articles as gold standards and reference concentration as the accuracy measure. Direct citation concentrates references more effectively than bibliographic coupling or co-citation, while journal-based discipline taxonomies are less accurate than topic-level document taxonomies.

  • Problem

    Taxonomy accuracy is difficult to determine despite its importance for citation-based evaluation and research-allocation decisions.

  • Method

    The study compares topic-level taxonomies built with direct citation, bibliographic coupling, and co-citation, using articles with at least 100 references as gold standards.

  • Results

    Direct citation concentrates references more than bibliographic coupling or co-citation and provides the most accurate topic-level taxonomy; journal-based discipline taxonomies are less accurate.

  • Takeaways & Limitations

    The authors recommend using the most accurate and comprehensive taxonomy methods rather than inferior methods based on tradition or convenience.

  • Takeaways & Limitations

    Accuracy is partly circular because the measure uses direct citations from the gold-standard articles, and the appropriate taxonomy granularity depends on the question.

Abstract

from arXiv · show

In 1965, Derek de Solla Price foresaw the day when a citation-based taxonomy of science and technology would be delineated and correspondingly used for science policy. A taxonomy needs to be comprehensive and accurate if it is to be useful for policy making, especially now that policy makers are utilizing citation-based indicators to evaluate people, institutions and laboratories. Determining the accuracy of a taxonomy, however, remains a challenge. Previous work on the accuracy of partition solutions is sparse, and the results of those studies, while useful, have not been definitive. In this study we compare the accuracies of topic-level taxonomies based on the clustering of documents using direct citation, bibliographic coupling, and co-citation. Using a set of new gold standards - articles with at least 100 references - we find that direct citation is better at concentrating references than either bibliographic coupling or co-citation. Using the assumption that higher concentrations of references denote more accurate clusters, direct citation thus provides a more accurate representation of the taxonomy of scientific and technical knowledge than either bibliographic coupling or co-citation. We also find that discipline-level taxonomies based on journal schema are highly inaccurate compared to topic-level taxonomies, and recommend against their use.

Early citationists

Early citation analysis pursued two related goals: tracing emerging research fronts and building taxonomic structures of scientific knowledge. Garfield, Kessler, Price, and Kuhn shaped this agenda by emphasizing citation linkages as tools for following scientific developments.

  • Early citationists: Garfield introduced citation analysis as a way to track the impact of seminal papers when subject indexes could miss emerging fields.His example followed Selye’s work through citing articles.
  • Early citationists: Garfield and colleagues used direct citations to trace the historical development of the DNA-code breakthrough.Their analysis found direct citation links connecting 65% of the selected papers, with additional links through interim papers.
  • Early citationists: Kessler used bibliographic coupling to show the contemporaneous structure of physics research, whereas Price viewed direct citation as taxonomic.The distinction was between a snapshot of current research and a systematic structure of published items.
  • Early citationists: Citation analysis was initially motivated by detecting emerging scientific communities and breakthroughs as they formed.Garfield, Kuhn, and related early work treated citation linkages as signals of emerging scientific developments.

Constrained progress

Global taxonomies expanded as computational barriers eroded, but early researchers chose co-citation and thresholding partly because direct-citation calculations were too demanding. Large-scale clustering later became feasible and reopened the question of which method best supports taxonomic mapping.

  • Constrained progress: Computational capacity constrained early attempts to create comprehensive taxonomies covering all of science.The relationship between taxonomy size and publication year suggests a persistent technical barrier tied to clustering and computer memory.
  • Constrained progress: Co-citation became the first method used for an all-science taxonomy because the 1972 dataset was too large for available direct-citation computation.That dataset contained 93,800 source documents and 867,600 references.
  • Constrained progress: Thresholding was chosen to isolate emerging topics, which represented only a small fraction of the literature at any given time.Early citationists assumed direct citation could identify emerging topics, but needed to reduce the data to make all-science clustering feasible.
  • Constrained progress: Single-link clustering created over-aggregated co-citation clusters through chaining effects.Small imposed a maximum cluster size, while Simon later developed a procedure to split large clusters into meaningful groups.
  • Constrained progress: Modern modularity-based and force-directed algorithms enabled clustering of millions of documents and supported large direct-citation taxonomies.Waltman and van Eck classified 10 million papers using 97 million direct citation links.
  • Constrained progress: Once computational barriers were reduced, the central unresolved issue became which citation strategy creates the most accurate taxonomy.A prior study combined direct citation for taxonomic structure with co-citation for detecting emerging areas, but did not identify the best citation-based taxonomy method.

Accuracy and gold standards

Accurately evaluating literature taxonomies is difficult because no physical ground truth exists and legitimate classifications can differ. This study therefore extends earlier objective comparisons by evaluating historical taxonomies with a new reference-concentration standard.

  • Accuracy and gold standards: Topic taxonomy accuracy lacks a physical ground truth, replicated measurements, or universally accepted gold standard.Consequently, many studies have relied on face validity rather than a defensible objective measure.
  • Accuracy and gold standards: Earlier accuracy studies commonly compared similarity measures or taxonomy outputs with established schemes such as Web of Science subject categories.These comparisons provided useful reference points but did not fully resolve the accuracy problem.
  • Accuracy and gold standards: Earlier work compared 13 methods for research-front clustering using concentration of PubMed grant-to-article linkages as an accuracy metric.The metric assumed that articles acknowledging the same grant should be concentrated in the same cluster.
  • Accuracy and gold standards: This study compares nine global document-based and seven global journal-based taxonomies to assess historical knowledge structures.The authors note that absolute accuracy is unattainable, but accurate perceived structure is needed for decisions by researchers, administrators, and funders.

Citation-based taxonomies

The study constructs comparable taxonomies from direct citation, bibliographic coupling, and co-citation using a common clustering approach and different temporal windows. It then evaluates them against reference-concentration gold standards to identify the most accurate citation-based representation.

  • Citation-based taxonomies: The study generates taxonomies from direct citation, co-citation, and bibliographic coupling for direct comparison.All nine document-based taxonomies use the CWTS modularity-based method, which accepts link strengths from any of these processes.
  • Citation-based taxonomies: Direct citation uses Scopus documents from 1996–2012, reflecting the method’s stronger performance with long time windows.The calculation also includes non-indexed documents cited at least twice.
  • Citation-based taxonomies: Bibliographic coupling uses a single-year 2010 window to represent which cited documents were considered important in current science.The authors characterize this short-window approach as largely obliterating historical information.
  • Citation-based taxonomies: Co-citation uses a three-year 2011–2013 window intended to produce a more stable solution than earlier single-year calculations.The expanded window became feasible because the SLM algorithm and Amazon Compute Cloud could handle hundreds of millions of edges.

Journal-disciplinary taxonomies

The study evaluates seven journal-based classification schemes and journal IDs as alternative discipline-level taxonomies alongside document-level taxonomies. These schemes differ in coverage, category assignment, and treatment of multidisciplinary journals.

  • Journal-disciplinary taxonomies: Seven journal classification schemes were linked to Scopus data through journal names for comparison with document-level taxonomies.The schemes include ASJC, UCSD, Science-Metrix, ARC, ECOOM, Web of Science categories, and NSF.
  • Journal-disciplinary taxonomies: Journal schemes differ in whether journals receive single or multiple category assignments.UCSD, Science-Metrix, and NSF use single categories, whereas ASJC, ARC, ECOOM, and Web of Science assign some journals to multiple categories.
  • Journal-disciplinary taxonomies: The evaluation also treats each journal or conference as its own knowledge category using Scopus journal IDs.This approach is included because journal-level categories are used in studies of innovativeness and science mapping.
  • Journal-disciplinary taxonomies: Table 3 reports matched source titles, category counts, and paper coverage for the seven journal-based partitions.Coverage is measured relative to papers from the DC5 document-level taxonomy.
  • Journal-disciplinary taxonomies: ASJC and journal IDs have the highest coverage, but each reaches only 60% because most DC5 non-source items lack journal IDs.UCSD has the second-highest coverage, while NSF has the lowest because its journal list excludes newer prominent journals.

Proposed gold standards and accuracy

The study addresses the difficulty of validating historical knowledge taxonomies by proposing long-bibliography papers as gold standards and measuring reference concentration across taxonomy clusters. It evaluates 2010 articles and reviews with at least 100 references, using the Herfindahl index while accounting for incomplete coverage.

  • Proposed gold standards and accuracy: Few studies rigorously measure taxonomy accuracy because literature partitions lack agreed-upon ground truth or gold standards.The authors therefore seek an accuracy measure suited to historical taxonomic subjects rather than current research fronts.
  • Proposed gold standards and accuracy: Papers with at least 100 references are proposed as gold standards because their reference concentration can evaluate competing taxonomies.This choice is supported by the role of synthesis and review papers in connecting current research with its historical development.
  • Proposed gold standards and accuracy: 3.14% of 2010 articles and reviews had at least 100 references, approximately one synthesis paper for every 32 documents.The study used reference-count bins to examine whether these documents behave like synthesis articles.
  • Proposed gold standards and accuracy: For papers with at least 100 references, the correlation between references and future citations drops to .0793 from .4034 for papers with fewer than 100 references.At this threshold, elite-author involvement has a stronger future-citation correlation, increasing from .2677 to .3887.
  • Proposed gold standards and accuracy: The final gold-standard set contains 37,207 2010 articles and reviews with at least 100 references and at least 80% of references available in DC5.Half are independently coded as reviews, and the set represents roughly one synthesis paper for every 35 articles.
  • Proposed gold standards and accuracy: The 37,207 gold-standard papers contribute 5,334,016 references spanning 79,354 of 91,726 DC5 clusters.Their references point to 3,248,243 unique papers, with 71.2% of union-set reference papers cited by only one gold-standard article.
  • Proposed gold standards and accuracy: Accuracy is compared with a Herfindahl index measuring reference concentration across clusters within each taxonomy.DC5 is the baseline because it contains 95.7% of gold-standard references; unavailable references are assigned their own cluster in every taxonomy.

Taxonomy-level results

Direct citation produces the most accurate topic-level taxonomy among the citation-based methods studied, while journal-based taxonomies are substantially less accurate than document-based ones.

  • DC taxonomies are more concentrated than BC or CC taxonomies at comparable cluster counts, indicating the most accurate document-level representation.The comparison should use taxonomies with roughly the same number of clusters because concentration depends strongly on cluster count.
  • 86.1% of gold standard papers have the highest Herfindahl index in DC5, compared with 13.7% for BC5 and very few for CC5.
  • CC5 provides the least concentrated and thus least accurate solution among taxonomies with roughly 105 clusters.
  • Journal-based taxonomies are much less accurate than document-based taxonomies, with JID slightly below even CC5 in concentration.Among journal schemas, SM performs best and ASJC worst; neither vendor-provided system matches some outside-researcher classifications.
  • DC solutions are less skewed than BC solutions at every level, so their higher Herfindahl values are not caused by more uneven cluster-size distributions.The reported skewness values are DC2-5: 1.42, 1.67, 1.65, 2.45 and BC2-5: 2.39, 4.20, 5.04, 3.89.
  • Overall, document-based taxonomies outperform journal-based taxonomies, and direct citation yields the most accurate topic-level taxonomy.

Detailed example

The Rafols2010 example shows how DC5 concentrates related references into coherent topic clusters, while cluster granularity remains a question that depends on the policy application.

  • Detailed example: Rafols2010 contains 103 references, 92 available in DC5, and is located in cluster DC5-9524.
  • Detailed example: DC and BC place many references in the top-ranked cluster, whereas CC5 has no dominant cluster and correspondingly low Herfindahl concentration.BC Herfindahl values are lower than DC values because of lower coverage; DC and BC also show hierarchical changes across levels.
  • Detailed example: The appropriate granularity likely depends on the specific policy question, and the study does not address whether topic or specialty level is preferable.
  • Detailed example: The top two DC5 clusters contain 44 and 8 Rafols2010 references, respectively, with DC5-9524 focused on science mapping and DC5-432 on impact and metrics.
  • Detailed example: Rafols2010’s subject-category map and overlay technique build directly on historical co-occurrence, knowledge-typology, and visualization developments represented in DC5-9524.
  • Detailed example: DC5-9524 produces about 100 papers annually, but Rafols2010 is its only paper with over 100 references from 1996–2012.The authors suggest, without establishing, that methodology-focused clusters may contain fewer large-reference papers than basic-science clusters.

Summary and Implications

The study finds that direct-citation, document-based taxonomies represent scientific and technical knowledge more accurately than bibliographic-coupling or co-citation taxonomies. These findings support using accurate taxonomies in research evaluation and allocation while cautioning against journal-based disciplinary classifications.

  • Research planning and evaluation require an accurate picture of science and technology’s structure for decisions across government, industry, academia, and nonprofits.
  • Direct citation concentrates references more highly than bibliographic coupling or co-citation, indicating more accurate topic-level clusters overall.The comparison uses papers with at least 100 references as gold standards and treats higher reference concentration as greater accuracy.
  • The accuracy measure is partly circular because it uses direct citations from gold-standard papers, although those citations are defended as first-order relatedness indicators.The 37,207 gold-standard papers contributed only 0.9% of the links used in direct-citation clustering.
  • Direct-citation taxonomies preserve substantially more historical literature, whereas bibliographic coupling emphasizes recent references.The differing age distributions may help explain differences in measured accuracy.
  • Document-based taxonomies represent disciplines more accurately than journal-based taxonomies, leading the Leiden Ranking to replace journal subject categories with an algorithmic taxonomy.
  • An accurate direct-citation taxonomy may support innovation metrics, resource-allocation decisions, research strategy, and identification of topics where future jumps or bridges may occur.The authors report a strong correlation between a topic-level innovation metric and STAR METRICS funding data, while noting that additional study is needed.
Loading 1511.05078v2…