Source-linked AI summary
Classification methods applied to credit scoring: A systematic review and overall comparison
Francisco Louzada, Anderson Ara, Guilherme B. Fernandes
TL;DR
Credit-scoring research spans diverse binary classification methods, while customer credit-history data remain difficult to access. This paper systematically reviews 187 studies from 1992–2015 and finds shifting technique use, including growing prominence of combined methods in recent years.
Problem
Credit-scoring research uses diverse classification methods, but confidential customer credit-history data remain difficult for researchers to access.
Method
The paper systematically reviews 187 credit-scoring studies from 1992–2015 using defined selection criteria and 12 analytical questions, supplemented by simulations across nine methodologies.
Results
Combined techniques became the most-used method in the latest period, accounting for 21.2% of reviewed studies, while genetic, fuzzy, and discriminant methods declined.
Takeaways & Limitations
The review documents changing methodological preferences and the continuing importance of major classification techniques in credit-rating research.
Takeaways & Limitations
The review includes only English-language journal papers indexed in four databases and excludes unpublished work, theses, books, conference proceedings, and white papers.
Abstract
from arXiv · showhide
The need for controlling and effectively managing credit risk has led financial institutions to excel in improving techniques designed for this purpose, resulting in the development of various quantitative models by financial institutions and consulting companies. Hence, the growing number of academic studies about credit scoring shows a variety of classification methods applied to discriminate good and bad borrowers. This paper, therefore, aims to present a systematic literature review relating theory and application of binary classification techniques for credit scoring financial analysis. The general results show the use and importance of the main techniques for credit rating, as well as some of the scientific paradigm changes throughout the years.
1. Introduction
The introduction frames credit scoring as a quantitative approach for assessing borrower default risk and classifying applicants, whose use expanded with modern statistical techniques and banking regulation. The paper addresses gaps in prior reviews through a systematic analysis of binary classification techniques applied to credit scoring from 1992 to 2015.
- Research gap: Prior literature reviews examined important or selected classification methods, while Lessmann et al. did not cover general genetic and fuzzy methodologies.The introduction identifies incomplete methodological coverage as a limitation of earlier reviews.
- Contribution: This paper presents a general systematic review of binary classification techniques for credit scoring, covering more than 20 years of research from 1992–2015 and 187 papers.The review aims to clarify practical credit-rating applications and changes over time.
2. Survey methodology
The study uses a systematic literature review to classify published research on credit-scoring techniques through defined search procedures, eligibility criteria, and four analytical categories. From 437 potentially eligible papers, 187 were included and assessed using 12 conceptual questions.
- Review design: The review defines sources, search procedures, and four classification categories: publication year, journal title, co-authors, and a conceptual scheme based on 12 questions.These categories were designed to understand the historical application of credit-scoring techniques.
- Search scope: The study searched ScienceDirect, Engineering Information, Reaxys, and Scopus, covering 20,500 titles from 5,000 publishers worldwide.The review was limited to published literature available in these databases.
- Eligibility criteria: Eligibility was restricted to English journal papers using “credit scoring” with related machine-learning, data-mining, classification, or statistical topics.Unpublished papers, dissertations, books, conference proceedings, white papers, and other publication forms were excluded.
- Selection procedure: 187 papers were included after 250 of 437 potentially related credit-scoring papers were discarded for failing the second selection criterion.The included papers were subjected to systematic review according to 12 questions about their conceptual and methodological scenarios.
- Analytical classification: The review groups papers by seven main objectives: proposing new rating methods, comparing traditional techniques, conceptual discussion, feature selection, literature review, performance-measure studies, and other issues.These objective categories organize papers with different specific aims into generally similar groups.
3. The main classification methods in credit scoring
This section introduces the main classification techniques used in credit scoring, including neural networks, support vector machines, regression models, decision trees, fuzzy logic, and genetic programming. It also describes feature selection, missing-value imputation, and confusion-matrix metrics as methodological components of credit-rating analysis.
- Neural networks: Neural networks combine explanatory variables through linear and nonlinear interactions across hidden layers to produce response variables.Applications include mixture-of-experts, radial basis function, and hybrid discriminant-analysis models.
- Support vector machine: Support vector machines classify observations by finding an optimal hyperplane that maximizes the geometric distance between binary categories.The method can use linear, polynomial, Gaussian, or sigmoidal separations.
- Linear regression: Linear regression relates borrower characteristics to a binary target, but its unbounded output cannot be interpreted as a probability.Ordinary least squares estimates the coefficient vector, and the conditional expectation may segregate good and bad borrowers.
- Decision trees: Decision trees construct tree-like decision rules from historical data to classify cases through if-then logical conditions.The usual algorithms are CHAID, CART, and C5.
- Logistic regression: Logistic regression estimates a linear combination of explanatory variables and the logit transformation of a binary response, yielding category probabilities.It is a traditional method often compared with other techniques or used in combinations, with regularized and limited variants also possible.
- Supporting methods: Feature selection improves classification by discarding irrelevant variables, while missing-value imputation addresses incomplete credit-analysis records.Confusion-matrix metrics compare model predictions with true response values and identify misclassifications.
4. Results and discussion
The review finds a growing credit-scoring literature, concentrated in selected journals and authors, with new-method propositions dominating objectives and classification practices shifting across four periods. Private datasets, limited missing-data imputation, confusion-matrix costs, and established comparison techniques characterize the reviewed studies.
- Publication year: Published papers grew rapidly after 2000, averaging 7.8 papers annually with a standard deviation of 7.6 from 1992–2015.Annual publication counts ranged from 0 to 25 papers.
- Scientific journals: The 187 papers appeared across 73 journals, led by Expert Systems with Applications at 27.81% and Journal of the Operational Research Society at 10.70%.Most papers concerned computer science, decision sciences, engineering, and mathematics journals.
- Main objectives: Proposing new credit-scoring methods was the dominant objective, covering 51.3% of papers, while hybrid methods led technique frequencies at almost 20%.Combined methods accounted for almost 15%, and support vector machines and neural networks each represented around 13%.
- Main classification techniques: Combined techniques became the most used method in period IV at 21.2%, while logistic regression reached 15.2% in recent years and matched neural networks in that period.Support vector machines peaked at 21.4% during period II; genetic, fuzzy, and discriminant-analysis methods declined.
5. Is there a better method? A comparison study
The comparison study evaluated the presented classification methods across three credit datasets using Approximate Correlation and F1-score under repeated handout validation. FUZZY and SVM generally showed the strongest predictive performance, while SVM combined high performance with low computational effort.
- Comparison framework: The study compared all presented methods using AC and FM across Australian, German, and Japanese Credit benchmark datasets.Each dataset used 1000 replications with 70% training and 30% test samples.
- Predictive performance: For p = 0.5, FUZZY achieved the greatest predictive performance, with SVM as the second method.The passage reports this result in the presence of imbalance in bad payers.
- Predictive performance: For p = 0.1, SVM achieved the greatest predictive performance, with FUZZY as the second method.TREES was often third best and independent of the imbalance, whereas NN most often lost predictive performance under imbalance.
- Overall comparison: Overall, SVM stood out for combining high predictive performance with low computational effort relative to the other analyzed methods.This conclusion follows the reported predictive-performance rankings and computational-time comparison.
- Computational effort: SVM required 0.37s per replication compared with 48.92s for FUZZY, despite both being among the methods with greater predictive performance.GENETIC and FUZZY had the highest computational effort.
6. Final comments
This systematic review analyzed 187 credit-scoring papers published from 1992–2015, finding a growing and significant research area with broadly similar predictive performance across methods. Neural networks, support vector machines, hybrid and combined techniques were prevalent, while validation practices and publication coverage remained important limitations.
- 187 credit-scoring papers published in scientific journals from 1992–2015 were systematically analyzed, revealing an increasing and significant research area.
- Hybrid techniques were especially common for proposing new credit-scoring rating methods, despite broadly similar predictive performance across methods.
- Neural networks, support vector machines, hybrid and combined techniques were the most common tools, while support vector machines showed high predictive performance and low computational effort.
- K-fold cross-validation and holdout were the most common validation methods, but their results require careful interpretation because random database distribution introduces subjectivity.
- Confusion-matrix metrics became more common over time, while few studies addressed missing data and many applied feature selection as preprocessing.
- The review was limited to English-language journal papers indexed in ScienceDirect, Engineering Information, Reaxys and Scopus, leaving other databases and publication types for future inclusion.
Appendix · Table A.1
Appendix Table A.1 presents the questions and possible responses used in the proposed systematic review.
- Appendix: Table A.1 lists the questions guiding the proposed systematic review.The table is identified as a list of review questions.
- Appendix: Table A.1 also lists the possible responses to those review questions.The table pairs the questions with possible responses.
- Appendix: The appendix includes Table A.1 as a review-design reference.The supplied passage identifies the material as an appendix table.
- Table A.1: The table concerns the proposed systematic review.Its title explicitly links the questions and responses to the proposed review.
- Table A.1: The table’s content is organized around questions and responses.These are the two content categories named in the table title.
- Table A.1: Table A.1 is titled “List of questions and of possible responses.”This title summarizes the table’s stated purpose and contents.
Table A.2
Table A.2 indexes 187 reviewed papers published from 1992 to 2015 using coded entries across questions Q01–Q12. The table lists study-specific codes for papers spanning 2008–2015.
- Table A.2: 187 reviewed papers are indexed for the 1992–2015 period.The table organizes the review corpus by publication and coded attributes.
- Table A.2: The index uses twelve fields labeled Q01 through Q12.The repeated table headers define the coded question columns used for each paper.
- Table A.2: Studies published from 2008 to 2012 are represented with varied letter and numeric codes across the indexed fields.Examples include Tsai (2008), Finlay (2009), and Brown and Mues (2012), each shown with study-specific entries.
- Table A.2: The index continues through studies published in 2013–2015, including papers with differing numbers of coded entries.Examples include Zhu and Hu (2013), Zhang et al. (2014), and Lessmann et al. (2015).