Source-linked AI summary
Coauthorship and Citation Networks for Statisticians
Pengsheng Ji, Jiashun Jin
TL;DR
The paper addresses the lack of prior study of statisticians’ coauthorship and citation networks. It constructs network datasets from papers in four leading statistical journals and analyzes centrality, finding that highly cited papers are concentrated in regularization methods.
Problem
Coauthorship and citation networks for statisticians had not yet been studied, despite statisticians’ partial knowledge of their community providing useful interpretive context.
Method
The authors compile coauthorship and citation datasets from papers published between 2003 and the first half of 2012 in four leading statistical journals, then analyze network centrality using degree, closeness, and betweenness measures.
Results
The 30 most cited papers account for 16% of total citation counts, with the most highly cited papers focused on regularization methods such as adaptive lasso and group lasso.
Takeaways & Limitations
The datasets support further research on social-network models, methods and theory, and statisticians’ research habits, patterns, and network structures.
Takeaways & Limitations
Results may differ if the study includes journals serving statisticians from other regions or different research focuses, such as Bioinformatics.
Abstract
from arXiv · showhide
We have collected and cleaned two network data sets: Coauthorship and Citation networks for statisticians. The data sets are based on all research papers published in four of the top journals in statistics from $2003$ to the first half of $2012$. We analyze the data sets from many different perspectives, focusing on (a) centrality, (b) community structures, and (c) productivity, patterns and trends. For (a), we have identified the most prolific/collaborative/highly cited authors. We have also identified a handful of "hot" papers, suggesting "Variable Selection" as one of the "hot" areas. For (b), we have identified about $15$ meaningful communities or research groups, including large-size ones such as "Spatial Statistics", "Large-Scale Multiple Testing", "Variable Selection" as well as small-size ones such as "Dimensional Reduction", "Objective Bayes", "Quantile Regression", and "Theoretical Machine Learning". For (c), we find that over the 10-year period, both the average number of papers per author and the fraction of self citations have been decreasing, but the proportion of distant citations has been increasing. These suggest that the statistics community has become increasingly more collaborative, competitive, and globalized. Our findings shed light on research habits, trends, and topological patterns of statisticians. The data sets provide a fertile ground for future researches on or related to social networks of statisticians.
1. Introduction.
The paper introduces coauthorship and citation network data for statisticians and analyzes centrality, communities, and evolving research patterns. It identifies key authors, research areas, and community structures while documenting changes in collaboration and citation behavior.
- Data and scope: The study collects coauthorship and citation networks from all papers published between 2003 and the first half of 2012 in four leading statistical journals.The journals are Annals of Statistics, Biometrika, JASA, and JRSS-B.
- Research relevance: The data sets support analysis of statisticians’ research habits and network structures and provide a basis for broader studies covering more journals and longer periods.The authors note that results may differ for journals serving other regions or disciplinary focuses.
- Centrality: Centrality analysis identifies prolific, collaborative, and highly cited authors, including Peter Hall, Raymond Carroll, Jianqing Fan, Joseph Ibrahim, and Hui Zou.The study uses degree, closeness, and betweenness centrality, with degree interpreted according to network type.
- Centrality: Among 14 hot papers, 10 concern Variable Selection, suggesting it as a hot research area alongside Covariance Estimation, Empirical Bayes, and Large-scale Multiple Testing.
- Community detection: Community analysis identifies meaningful groups including High Dimensional Data Analysis, Theoretical Machine Learning, Dimension Reduction, Objective Bayes, Large-Scale Multiple Testing, Variable Selection, and Spatial and semi-parametric/nonparametric Statistics.The coauthorship network is fragmented into many disconnected components, while SCORE and D-SCORE identify communities in coauthorship and citation networks.
- Productivity, patterns and trends: Over 2003–2012, papers per author and self-citations decrease while distant citations increase, suggesting greater collaboration, competition, and globalization.
2. Centrality.
The paper uses multiple centrality measures to identify prominent authors and highly connected papers in statistician networks. These measures consistently highlight a small group of leading authors and variable selection as a particularly prominent research area, while citation counts remain limited by the selected data set.
- Author centrality: Different centrality measures largely agree in identifying Raymond Carroll, Jianqing Fan, and Peter Hall as the top three authors.The comparison covers degree, closeness, and betweenness centrality across the author-paper, coauthorship, and author-citation networks.
- Hot papers and areas: The 14 identified hot papers are concentrated especially in variable selection, with other possible hot areas including covariance estimation, empirical Bayes, and large-scale multiple testing.The hot papers were identified using degree, closeness, and betweenness centrality.
- Hot papers and areas: The three most cited papers received 75, 64, and 49 citations, respectively, and all concern high-dimensional variable selection.They address adaptive lasso, graphical lasso, and the Dantzig Selector.
- Citation patterns: The 30 most cited papers account for 16% of all citation counts, and the most highly cited papers mainly concern regularization methods.Examples include adaptive lasso and group lasso.
- Citation patterns: Citation-based prominence is bounded by the data set: Efron et al. (2004) has 4,900 Google Scholar citations but only 11 citations in these networks.By comparison, the adaptive lasso paper has 75 citations in the data set.
- Interpretation: The authors report centrality measures descriptively and do not intend them to rank one author or research area above another.The measures are described as natural choices or existing measures.
3. Community detection for Coauthorship networks.
The paper compares community-detection methods across two coauthorship networks, finding interpretable research groups but substantial method-dependent disagreement, especially in the more connected network.
- Coauthorship network (A): Coauthorship network (A) contains 2985 components among 3607 authors; 2805 components are singletons, 105 are pairs, and the average component size is 1.2.
- Coauthorship network (A): The giant 236-node component contains two communities: North Carolina and Carroll-Hall, although the Fan group is assigned differently across methods.
- Coauthorship network (A): The Fan group’s strong ties to both communities may indicate three communities, but four-method results for K = 3 are inconsistent.
- Coauthorship network (A): Theoretical Machine Learning and Dimension Reduction form the next two largest components, with 15 and 14 nodes, respectively.
- Coauthorship network (A): The five subsequent components correspond to Johns Hopkins, Duke, Stanford, Quantile Regression, and Experimental Design research groups.
- Coauthorship network (B): For Coauthorship network (B), SCORE, NSC, and APL identify Objective Bayes, Biostatistics, and High Dimensional Data Analysis communities, while BCPL disagrees substantially.
- Coauthorship network (B): SCORE and NSC differ over about 200 biostatisticians, with NSC assigning them to HDDA-Coau-B and SCORE to Biostat-Coau-B.
- Coauthorship network (B): APL may underestimate Objective Bayes and Biostat-Coau-B communities, whose reported sizes vary across methods from 20 to 69 and from 50 to 388, respectively.
4. Community detection for Citation network.
The Citation network is analyzed on its weakly connected giant component using directed community-detection methods, especially D-SCORE, which identifies three interpretable communities. D-SCORE combines directed spectral information with citer and citee network structure, while comparisons reveal method disagreement in some settings.
- Methods: D-SCORE adapts SCORE to directed networks by using left and right singular vectors and constructing citer and citee networks.Citer edges connect authors with a common citee, while citee edges connect authors with a common citer.
- Methods: D-SCORE clusters nodes across four subsets formed from the citer and citee giant components, assigning nodes through combined spectral coordinates, community centers, or weak-edge counts.The final subset is assumed small; in these data, it contains 14 nodes.
- Results: For K = 3, similar patterns in two D-SCORE panels support three Citation-network communities.The corresponding figure illustrates the method on the statistical citation network.
- Network and setup: The Citation network contains 3,607 authors; its weakly connected giant component has 2,654 authors, or 74% of all nodes.The analysis restricts attention to this giant component, which contains 14 nodes outside the citer and citee giant components.
D R Cox
Citation-network communities provide additional insight into coauthorship communities and reveal cross-network structures. The comparisons connect large coauthorship groups to citation-defined topics while also exposing smaller, geographically or institutionally linked groups.
- Comparison with Coauthorship network (A): The Citation network’s giant component contains 236 authors, with nearly all but 3 assigned to 3 D-SCORE communities.The communities include 60 authors in Spatial Statistics and Semi-parametric/Non-parametric statistics, 166 in Variable Selection, and 7 in Large-Scale Multiple Testing.
- Comparison with Coauthorship network (A): The Carroll-Hall group is strongly connected to Variable Selection, while the North Carolina group is strongly connected to Biostatistics.Raymond Carroll has close ties to both groups, receiving different assignments in coauthorship and citation analyses.
- Comparison with Coauthorship network (A): Theoretical Machine Learning, Dimension Reduction, Duke, and Quantile Regression are almost subsets of Variable Selection, while Stanford is almost a subset of Large-Scale Multiple Testing.Johns Hopkins is almost a subset of Spatial Statistics; Experimental Design spreads across areas.
- Comparison with Coauthorship network (B): In Coauthorship network (B), Objective Bayes separates into a 55% part tied to James Berger and a 25% part assigned to Variable Selection.The two parts therefore connect distinct coauthorship patterns to different citation-defined research areas.
- Comparison with Coauthorship network (B): High Dimensional Data Analysis splits into parts labeled as Spatial and Semi-parametric/Non-parametric Statistics, Variable Selection, and Large-Scale Multiple Testing.The authors associate these parts with corresponding topical leadership patterns, including spatial or biostatistical interests, variable selection, and FDR control.
- Comparison with Coauthorship network (B): Large-Scale Multiple Testing includes a 221-node subset of High Dimensional Data Analysis, a 115-node group outside its giant component, and a 17-node Bioinformatics group.The 115-node group includes many German researchers tied to Helmut Finner, while the Bioinformatics researchers publish relatively few papers in the four journals during the study period.
5. Discussions.
The paper identifies about 15 meaningful statistical communities while noting important scope, methodological, and informational limitations. The datasets remain useful for studying U.S.-based statisticians and for motivating broader future work.
- Community structures: About 15 meaningful communities include Spatial Statistics, Dimension Reduction, Large-Scale Multiple Testing, Objective Bayes, Quantile Regression, Theoretical Machine Learning, and Variable Selection.
- Limitations: The datasets cover papers from only four core statistical journals during 2003–2012, excluding many statisticians and publications in other disciplines.
- Conclusions: The datasets and analyses support understanding networks of statisticians whose home base is the USA and provide a starting point for more complete publication data.
- Limitations: Community labels do not always accurately represent all authors or papers because communities can be difficult to interpret.
- Future research: The paper leaves mixed membership, link prediction, recognition, and distinctions among important, influential, and popular work for future research.
- Limitations: Coauthorship and citation networks provide limited information about research habits and trends compared with abstracts, affiliations, keywords, or full papers.
6. Appendix I: Productivity, patterns and trends.
The appendix examines productivity, collaboration, and citation patterns in the 2003–2012 data. It finds increasingly collaborative and globally connected behavior, alongside highly skewed publication and citation distributions.
- Productivity: 3,248 papers and 3,607 authors yield an overall average of 0.90 paper per author.
- Productivity: The top 10% most prolific authors contributed 41% of papers, while the top 20% contributed 58%.
- Productivity and collaboration: The coauthorship degree distribution and non-divided paper-count distribution have power-law tails, whereas divided contributions remain highly skewed without a power-law tail.
- Citation patterns: The top 10% of highly cited papers received about 60% of citations, and the top 20% received about 80%, with Gini coefficient 0.77.
- Citation patterns: Coauthor citations showed 79% reciprocation, compared with 25% for distant citations.
- Citation patterns and trends: Self-citations decreased, coauthor-citation proportions stayed roughly constant, and distant citations increased over the 10-year period.
- Citation patterns: The overall mean citation delay was 3.30 years, with delays of 2.81, 3.36, and 3.51 years for self-, coauthor, and distant citations.
7. Appendix II: Data collection and cleaning.
The authors assembled and cleaned a publication, authorship, and citation dataset from four journals using bibliographic sources, DOI matching, citation extraction, and author disambiguation. They released cleaned files and clustering rules to support reproducibility.
- Scope: The dataset covers papers in AoS, JASA, JRSS-B, and Biometrika from 2003 through the first half of 2012.
- Cleaning: After removing reviews, corrections, and other non-original items, the cleaned dataset contains 3,248 papers.
- Collection pipeline: Data collection identified papers, extracted citations among them, and identified all authors for each paper.
- Paper identification: DOIs were used as paper identifiers, with Web of Science and MathSciNet combined to address missing DOI records.
- Author identification: Author disambiguation addressed incomplete names, inconsistent name formats, and distinct authors sharing the same name.
- Reproducibility: The authors prepared raw and cleaned bibliographic files, clustering rules, and author lists for reproducibility.