Source-linked AI summary

A Bibliometric Review of Large Language Models Research from 2017 to 2023

Lizhou Fan, Lingyao Li, Zihui Ma, Sanggyu Lee, Huizi Yu, Libby Hemphill

arXiv:2304.02020v1cs.DLcs.CLcs.CYcs.SI

TL;DR

Existing LLM research was expanding rapidly but was often limited to specific tasks or applications, motivating a broader evidence-based view. The paper applies bibliometric and discourse analyses to over 5,000 publications from 2017 to early 2023, identifying research themes, applications, and collaboration patterns. It finds a rapidly evolving field spanning algorithms, NLP, and diverse domains, while emphasizing the value of historical understanding and continued monitoring.

  • Problem

    Prior LLM literature often focused on specific NLP tasks or applications, creating a need to understand broader research paradigms and collaborations as the field expanded.

  • Method

    The paper applies bibliometric and discourse analyses to Web of Science literature, using topic modeling for research paradigms and network analysis for scholarly collaborations.

  • Results

    The review covers over 5,000 publications and identifies a fast-evolving LLM research landscape spanning algorithms, NLP tasks, diverse applications, and international and organizational collaborations.

  • Takeaways & Limitations

    The review serves as a roadmap for researchers, practitioners, and policymakers navigating LLM research and identifying knowledge gaps and opportunities.

  • Takeaways & Limitations

    The study’s state-of-the-art discussion may be surpassed by newer advances because LLM research is evolving rapidly.

Abstract

from arXiv · show

Large language models (LLMs) are a class of language models that have demonstrated outstanding performance across a range of natural language processing (NLP) tasks and have become a highly sought-after research area, because of their ability to generate human-like language and their potential to revolutionize science and technology. In this study, we conduct bibliometric and discourse analyses of scholarly literature on LLMs. Synthesizing over 5,000 publications, this paper serves as a roadmap for researchers, practitioners, and policymakers to navigate the current landscape of LLMs research. We present the research trends from 2017 to early 2023, identifying patterns in research paradigms and collaborations. We start with analyzing the core algorithm developments and NLP tasks that are fundamental in LLMs research. We then investigate the applications of LLMs in various fields and domains including medicine, engineering, social science, and humanities. Our review also reveals the dynamic, fast-paced evolution of LLMs research. Overall, this paper offers valuable insights into the current state, impact, and potential of LLMs research and its applications.

1. Introduction

LLMs emerged as general-purpose language models that shifted NLP research and attracted broad study, but prior literature often focused on individual tasks or applications. This paper maps research paradigms, collaborations, and knowledge gaps across the expanding LLM literature.

  • LLM foundations: LLMs use transformer-based neural networks with billions of parameters trained on massive unlabeled text through self-supervised learning.Self-attention helps transformers capture contextual relationships and long-range dependencies between input tokens.
  • LLM foundations: Since their emergence in 2018, LLMs have demonstrated strong performance across diverse NLP tasks and shifted research toward general-purpose models.Examples include BERT, GPT families, and LLaMA.
  • Research gap: Earlier research emphasized specific NLP tasks and applications, leaving a need for a broader view of LLM research.Prior work covered areas such as medicine, health sciences, and politics but was often task- or application-specific.
  • Paper scope: The paper examines research paradigms through topic modeling and discourse analysis, spanning algorithms, NLP tasks, applications, infrastructures, and critical studies.It also analyzes scholarly collaboration networks from international and organizational perspectives.
  • Paper scope: The review provides an up-to-date bibliometric analysis and a roadmap for researchers, practitioners, and policymakers to identify knowledge gaps and opportunities.The authors position the analysis as support for navigating the current LLM research landscape.

2. Background

LLMs are pretrained, deep-learning language models whose scale and prompting capabilities support broad language tasks and multidisciplinary applications. Their development progressed from transformer architectures and pretrained models to increasingly large systems, alongside continuing computational demands.

  • Foundations: LLMs are pretrained deep-learning models trained and fine-tuned on vast text collections to learn language patterns and build a language knowledge base.Unlike conventional task-specific models, they can perform multiple tasks with few prompts.
  • History: Transformers addressed earlier difficulties with capturing long-range dependencies that limited recurrent and convolutional neural networks.This development supported later progress in translation, summarization, and question-answering.
  • History: BERT introduced bidirectional pretraining in 2018, while GPT-2 used a 1.5-billion-parameter transformer to generate diverse linguistic outputs.The passages note computational expense for both BERT pretraining and GPT-2 training and operation.
  • History: GPT-3, introduced in 2020, contained 175 billion parameters and generated high-quality text with little to no fine-tuning.Its development used a higher layer count and more diverse training data.
  • NLP tasks: LLM research covers fine-tuning and prompting across tasks including relation extraction, dialogue analysis, summarization, sentiment analysis, named entity recognition, and classification.The cited studies report potential improvements in NLP task accuracy and fluency.
  • Applications: Applications span medicine and engineering, including diagnostic assistance, treatment suggestions, medical education, document analysis, emergency planning, and building-defect detection.These examples illustrate use across professional, domain-specific workflows.

3. Data and Methods

The study combines Web of Science bibliometric data with topic modeling, discourse analysis, and collaboration-network analysis to characterize LLM research from 2017 to early 2023. It identifies themes from publication content and examines scholarly relationships across the field.

  • Workflow: The workflow collects LLM publication metadata, analyzes research paradigms through thematic discourse, and studies scholarly collaboration networks.The workflow is presented as the overall data-and-methods process.
  • Data: Web of Science searches combined LLM-related terms and model names across article titles and topic fields, covering 2017-01-01 through 2023-02-20.The search returned 5752 publications.
  • Topic modeling: Topic modeling discovers latent semantic topics by representing documents as combinations of topics characterized by word distributions.The method was used to analyze LLM research paradigms.
  • Topic modeling: BERTopic used Sentence-BERT embeddings of title–abstract combinations, and the analysis selected 200 clusters to balance topical similarity against excessive fragmentation.The study characterized the resulting topics and grouped them into five higher-level research themes.
  • Topic modeling: Two authors independently annotated all 200 topics, reaching Krippendorff’s α = 0.76 initially and full agreement after discussion.Tableau was used to visualize annotated results and the publication corpus.
  • Network analysis: CiteSpace generated co-citation and collaboration networks to identify significant publications and examine scholarly relationships among countries, institutions, and authors.Its co-citation clustering uses an expectation maximization algorithm that assigns each reference to one cluster.

4. Results

LLM research expanded sharply from 2017 to 2023 and spans algorithmic, applied, critical, and infrastructural themes. The results also show semantically connected subfields and increasingly international, cross-sector collaboration.

  • Research trends: Publication counts increased steadily from 2017 to 2023, with a sharp 2019–2020 spike and continued growth afterward.The spike is associated with interest in transformer-based NLP algorithms and the public release of GPT-3.
  • Research themes: 54% of publications, or 2,980 of 5,527, concerned Algorithm and NLP Tasks, making it the largest research theme.Social and Humanitarian Applications accounted for 25% (1,387 of 5,527), while Medical and Engineering Applications represented 18% (1,006/5,527).
  • Research themes: Algorithm and NLP Tasks and Social and Humanitarian Applications were distributed across the topic map, indicating semantic closeness and communication with other subdomains.Medical and Engineering Applications formed a professional cluster, while critical studies were semantically close to the subjects they analyzed.
  • Key discourses: Co-citation networks identified general NLP and machine-learning keywords centrally, task-specific keywords peripherally, and diverse application subdomains supported by pre-trained models and core NLP tasks.Social and humanitarian research prominently addressed fake news, Twitter, hate speech, rumor detection, and argumentation mining.
  • International collaborations: The USA, England, and India had the highest collaboration-network degree values at 51, 41, and 35, while China and the USA had the most publications at 1,828 and 1,344.The USA occupied a more centralized collaboration position than China despite China’s higher publication frequency.
  • Institutional collaborations: Universities dominated institutional collaboration networks, with 14 of the top 20 organizations being universities, while major technology companies increasingly participated in academic–industry partnerships.The Chinese Academy of Sciences ranked first with 205 papers, and recent projects included collaborations involving Microsoft, Google, Tencent, and university systems.

5. Discussion

The discussion highlights LLMs’ rapid, interdisciplinary expansion alongside persistent technical, ethical, and methodological constraints. It also emphasizes collaboration networks and the study’s role in guiding future research and policy.

  • Research evolution and applications: LLMs research has advanced rapidly, with applications spanning medical, engineering, social, and humanitarian domains.The paper attributes this expansion to developments in natural language processing and diverse research activity.
  • Research evolution and applications: Knowledge transfer occurs across LLMs subdomains, including between algorithmic, NLP-task, social, humanitarian, medical, and engineering research.The discussion links this transfer to semantic closeness, interdisciplinary work, and the adaptability of pretrained models to specialized tasks.
  • Challenges: Computational complexity, limited interpretability, and ethical concerns constrain the development and application of some LLMs.The paper connects these constraints to scalability, trust, adoption, environmental costs, bias, privacy, and malicious use.
  • Collaboration networks: The study presents collaboration-network findings as guidance for researchers, funders, policymakers, and nonprofit organizations seeking impactful engagement in LLMs research.It also discusses roles for government agencies, universities, companies, data creators, and infrastructure providers.
  • Collaboration networks: International and organizational collaboration networks are increasingly significant, with growing participation from countries and institutions through 2022.The analysis identifies prominent universities and companies and reports close academia–industry relationships involving computing resources, funding, testing, and validation.
  • Limitations and outlook: The bibliometric review is limited by possible keyword-based false inclusions, imperfect topic clustering, and incomplete Web of Science coverage.The authors describe human annotation, model comparisons, and the representativeness of the available collection for collaboration analysis.

6. Conclusion

The study surveys over 5,000 LLMs papers from 2017 to early 2023 using discourse and bibliometric analysis. It identifies rapid NLP advances, broad applications, collaboration-driven knowledge sharing, and continuing technical and ethical challenges.

  • Over 5,000 LLMs research papers from 2017 to early 2023 were analyzed through discourse and bibliometric methods.
  • LLMs research has produced significant NLP advances and diverse applications across medical, engineering, social, and humanitarian domains.
  • Interdisciplinary, inter-organizational, and international collaborations support knowledge sharing and the development of LLMs research.
  • Computational complexity, limited interpretability, and ethical concerns remain challenges in designing and applying LLMs.
  • The paper calls for greater openness and cooperation among stakeholders to support responsible LLM development and application.

A. Topic word scores

Figure 9 presents example topic word scores.

  • Figure 9 is titled “Example topic word scores.”
  • The figure concerns topic modeling within the paper’s analysis of LLMs research.
  • The supplied passage identifies Figure 9 as an example rather than reporting specific scores or comparisons.

B. Topic modeling and research themes

The topic analysis identifies broad LLM research themes spanning algorithms, NLP tasks, social and humanitarian applications, medical and engineering applications, critical studies, and infrastructure. The listed topics range from translation, speech, question answering, and sentiment analysis to clinical applications, cybersecurity, social media, and hardware optimization.

  • Social and Humanitarian Applications: Social and Humanitarian Applications cover fake-news detection, cyberbullying, pandemic tweets, mental health, financial forecasting, reviews, education, and gender profiling.
  • Medical and Engineering Applications: Medical and Engineering Applications include medical records, diagnosis, clinical named-entity recognition, biomedical advice, drug-related analysis, radiology reports, software defects, and engineering hazards.
  • Algorithm and NLP Tasks: Algorithm and NLP Tasks include translation, speech recognition, question answering, sentiment analysis, dialogue, parsing, entity recognition, embeddings, and knowledge graphs.
  • Critical Studies: The topic list also includes critical studies of GPT-3, artificial intelligence, gender bias, and sexism, alongside removed topics unrelated to the main LLM themes.
  • Infrastructure: Infrastructure topics address memory, parallelism, GPUs, distributed systems, quantization, hardware accelerators, pruning, inference, and model compression.
Loading 2304.02020v1…