Source-linked AI summary

Machine learning on small size samples: A synthetic knowledge synthesis

Peter Kokol, Marko Kokol, Sašo Zagoranski

arXiv:2103.01002v1cs.LGcs.AI

TL;DR

Machine learning on small-size datasets poses a problem despite various approaches to solve it. This study uses bibliometric and content analysis to examine the small-dataset research literature, finding 1,254 publications by 3,833 authors, rising publication activity, and high field maturity.

  • Problem

    Machine learning on small-size datasets poses a problem despite various approaches to solve it.

  • Method

    The study uses bibliometric and content analysis of small-dataset research, including higher-abstraction analysis into categories.

  • Results

    1,254 publications by 3,833 authors were identified, while publication activity rose over time and the field reached a high level of maturity.

  • Takeaways & Limitations

    Frequently reported difficulties include small dataset size, high or low dimensionality, and unbalanced data.

  • Takeaways & Limitations

    The analysis was limited to publications that included qualitative components, which might bias the study results.

Abstract

from arXiv · show

One of the increasingly important technologies dealing with the growing complexity of the digitalization of almost all human activities is Artificial intelligence, more precisely machine learning Despite the fact, that we live in a Big data world where almost everything is digitally stored, there are many real-world situations, where researchers are faced with small data samples. The present study aim is to answer the following research question namely What is the small data problem in machine learning and how it is solved?. Our bibliometric study showed a positive trend in the number of research publications concerning the use of small datasets and substantial growth of the research community dealing with the small dataset problem, indicating that the research field is moving toward higher maturity levels. Despite notable international cooperation, the regional concentration of research literature production in economically more developed countries was observed.

ITRODUCTION

Machine learning on small datasets is a significant research problem because smaller datasets generally reduce pattern-recognition power and accuracy. The study addresses this gap by synthesizing and structuring existing evidence on the small dataset problem, its solutions, and research gaps.

  • ITRODUCTION: Smaller datasets generally make machine-learning algorithms less powerful and less accurate at recognizing patterns.The paper links machine-learning pattern-recognition power to dataset size.
  • ITRODUCTION: The study addresses the lack of holistic research on the small dataset problem in machine learning.
  • ITRODUCTION: Synthetic knowledge synthesis aggregates current evidence to overcome isolated findings that may be incomplete or lack overlap with other solutions.
  • ITRODUCTION: The analysis extracts, synthesizes, and multidimensionally structures a potentially comprehensive corpus of scholarship on small datasets.
  • ITRODUCTION: The study identifies research gaps and maps important themes, solutions, and relationships for researchers, practitioners, and newcomers.

METHODOLOGY

The study uses synthetic knowledge synthesis to analyze publications on small datasets in machine learning. It combines Scopus harvesting, text mining, bibliometric mapping, clustering, thematic labeling, and research-dimension analysis.

  • METHODOLOGY: The research asks what the small data problem is and how it is solved.
  • METHODOLOGY: Synthetic knowledge synthesis harvests publications, codes their content, maps research clusters, analyzes connections, labels themes, and identifies research dimensions.
  • METHODOLOGY: The search used Scopus and a TITLE-ABS-KEY query covering small databases, datasets, or samples with machine learning while excluding large counterparts.
  • METHODOLOGY: The search was performed on 7 January 2021, and publication metadata were exported as a CSV-formatted corpus.
  • METHODOLOGY: VOSViewer performed bibliometric mapping and text mining, clustering closely associated author keywords into color-coded groups.

RESULTS AND DISCUSSION

The bibliometric analysis found rapid growth and increasing maturity in small-dataset machine-learning research, alongside strong international collaboration and concentration in economically developed countries. Research themes center on addressing small-data challenges through dimensionality reduction, data augmentation, and established machine-learning methods across diverse application domains.

  • Research growth and composition: 1,254 publications by 3,833 authors document a substantial research community studying machine learning on small datasets.The corpus included 500 conference papers, 33 review papers, 17 book chapters, and 17 other publication types.
  • Geographic and institutional distribution: China and the United States led national and institutional output, while the most productive countries were predominantly members of the G9 economically strong countries.China produced 439 publications and the United States 297; the Chinese Academy of Sciences was the most productive institution with 52 publications.
  • Research growth and composition: Publications began in 1995, rose linearly from 2003, and grew exponentially in 2016, with review-paper production becoming continuous that year.The authors assess the field as being in the third stage of Schneider’s scientific-discipline evolution model.
  • International collaboration: Among countries producing at least 10 papers, an elaborate international co-authorship network emerged, with the strongest cooperation between the United States and China.Russia, South Korea, and Brazil were identified as the newest countries joining the network.
  • Themes and solutions: Four prevailing themes emerged, led by dimension reduction in complex big-data analysis and data augmentation in deep learning.Frequently used approaches included support vector machines, decision trees or forests, convolutional neural networks, transfer learning, and principal component analysis.
  • Themes and solutions: The main reported small-data difficulties were limited dataset size, high or low dimensionality, and class imbalance across fields including bioinformatics, imaging, engineering, forecasting, and healthcare.The review also identifies support vector machines, decision trees or forests, convolutional neural networks, transfer learning, and principal component analysis among commonly used responses.

Strengths and limitations

The study’s main strength is its pioneering bibliometric and content analysis of small-dataset research, while its scope is constrained by Scopus-only coverage and possible qualitative bias.

  • The study presents the first bibliometric and content analysis of small dataset research.
  • The analysis was limited to publications indexed in Scopus.The authors describe Scopus as indexing the largest and most complete set of information titles.
  • The qualitative components of the analysis might bias the study’s results.

Conclusion

The bibliometric study finds increasing publication activity and research-community growth around small datasets, alongside international cooperation and regional concentration in economically developed countries. Its multidimensional landscape can support understanding and further knowledge development among researchers, practitioners, and other stakeholders.

  • The number of research publications concerning small datasets showed a positive trend.
  • Substantial research-community growth indicates that small-dataset machine learning is moving toward higher maturity levels.
  • Despite notable international cooperation, research literature production remained regionally concentrated in economically more developed countries.
  • The study provides a multidimensional view of the small-datasets problem and its scientific landscape.
  • These results can help researchers and practitioners improve their understanding and catalyse further knowledge development.
  • The study can inform novice researchers, interested readers, research managers, and evaluators.
Loading 2103.01002v1…