Source-linked AI summary

The Value of Big Data for Credit Scoring: Enhancing Financial Inclusion using Mobile Phone Data and Social Network Analytics

María Óskarsdóttir, Cristián Bravo, Carlos Sarraute, Jan Vanthienen, Bart Baesens

arXiv:2002.09931v1cs.SIcs.CYcs.LGstat.ML

TL;DR

This paper examines whether mobile phone and call-network data add value to credit scoring beyond existing techniques and traditional information. It introduces mobile phone data as a Big Data source and reports a decrease of up to 47.5% in credit availability from positive information.

  • Problem

    Existing techniques could offer only marginal performance gains, motivating the question of the added value of including call data in credit scoring.

  • Method

    The paper introduces mobile phone data as a Big Data source for credit scoring and uses call networks to propagate delinquent-customer influence through connected nodes.

  • Results

    Positive information leads to a decrease of up to 47.5% in credit availability.

  • Takeaways & Limitations

    Mobile phone data should be used strictly in a positive framework, while such insights could facilitate credit access for borrowers with little or no credit history.

  • Takeaways & Limitations

    Using mobile phone data to profile repayment behavior presents an ethical challenge because borrowers could be unfairly punished.

Abstract

from arXiv · show

Credit scoring is without a doubt one of the oldest applications of analytics. In recent years, a multitude of sophisticated classification techniques have been developed to improve the statistical performance of credit scoring models. Instead of focusing on the techniques themselves, this paper leverages alternative data sources to enhance both statistical and economic model performance. The study demonstrates how including call networks, in the context of positive credit information, as a new Big Data source has added value in terms of profit by applying a profit measure and profit-based feature selection. A unique combination of datasets, including call-detail records, credit and debit account information of customers is used to create scorecards for credit card applicants. Call-detail records are used to build call networks and advanced social network analytics techniques are applied to propagate influence from prior defaulters throughout the network to produce influence scores. The results show that combining call-detail records with traditional data in credit scoring models significantly increases their performance when measured in AUC. In terms of profit, the best model is the one built with only calling behavior features. In addition, the calling behavior features are the most predictive in other models, both in terms of statistical and economic performance. The results have an impact in terms of ethical use of call-detail records, regulatory implications, financial inclusion, as well as data sharing and privacy.

1. Introduction

The paper argues that improving credit scoring should prioritize innovative Big Data sources over increasingly sophisticated classifiers. It introduces mobile phone call data within a positive-information framework and evaluates its statistical, profit, financial-inclusion, ethical, and regulatory implications.

  • Traditional credit-scoring research found only marginal gains from newer classification techniques, motivating attention to innovative Big Data sources.
  • Call-detail records can represent borrowers’ social networks and calling behavior, but using them to assess repayment may be unfair to borrowers.
  • 47.5%: excluding positive information can reduce credit availability by up to 47.5%.
  • The paper introduces mobile phone data for credit scoring and argues that it should be used strictly as positive information to expand financing access.
  • The study quantifies mobile phone data’s value through both statistical performance, including AUC, and profit, using a combined banking, sociodemographic, and CDR dataset.
  • The dataset combines up to one and a half years of banking history with calling activity from almost 90 million unique phone numbers.
  • The research asks whether call data adds value, can replace traditional scoring data, and reveals how default behavior propagates through call networks.
  • The study evaluates these questions from statistical and profit perspectives and considers implications for financial inclusion, regulation, data sharing, and privacy.

2. Related Work

Related work shows that social connections and behavioral similarity can predict creditworthiness, but network construction, confounding, and empirical validation remain important challenges. Prior research increasingly considers call networks as an alternative credit-scoring data source.

  • Social network learning incorporates social behavior patterns and joint customer actions into predictive models.
  • Social-network effects have been reported in churn prediction and credit-card fraud detection, where networks connect related entities.
  • Default correlation is widely assumed in credit scoring, although regulatory correlation values have sometimes been set arbitrarily or through unpublished procedures.
  • A central challenge is defining the network, while observed social behavior may reflect homophily, social influence, or external confounding factors.
  • Online ties may not reveal true creditworthiness and can be manipulated, highlighting a limitation of social-network-based scoring.
  • Prior studies found that interacting friends can be more predictive than non-interacting friends, while implicit behavioral-similarity networks can outperform both.
  • Research on call networks links individuals through contact records and finds socioeconomic homophily, including connections among people sharing socioeconomic class.
  • Call-network credit-scoring proposals included theoretical work whose models lacked an important empirical evaluation.

3. Methodology

The methodology extracts information from call-detail records through social networks and influence propagation, then evaluates models and features using profit-based measures.

  • The proposed methodology extracts credit-scoring information from CDR data using social networks and influence propagation.
  • The paper presents techniques for evaluating model and feature performance in terms of profit.

3.1. Call Networks: Featurization and Propagation

The paper constructs call networks from phone-contact records, extracts direct and whole-network features, and propagates delinquency influence from known delinquent customers. Personalized PageRank and Spreading Activation produce indirect exposure features for credit-scoring models.

  • Call-network construction: A call network represents people appearing in CDR logs as nodes connected by phone-call edges.
  • Call-network construction: Edges may be undirected or directed, and weights can encode relationship intensity such as call count or duration.
  • Labels and features: Customers are labeled as defaulters or non-defaulters, while customers with one or two months of arrears are also treated as delinquent.
  • Influence propagation: Known delinquent customers serve as sources whose influence is propagated through social ties to create exposure scores.
  • Labels and features: Direct features summarize delinquent customers in a node’s first-order neighborhood, whereas indirect features use the whole network structure.
  • Influence propagation: Personalized PageRank and Spreading Activation model longer-range influence; both yield indirect network features for scoring models.
  • Link-based exposure features: Propagation-based exposure scores are thresholded to classify nodes as low- or high-risk, after which new link-based features are extracted.
  • Personalized PageRank: Personalized PageRank assigns higher exposure scores to nodes closer to delinquent source nodes.

3.2. The Expected Maximum Profit Measure

The Expected Maximum Profit measure evaluates credit-scoring models by expected loan losses and operational income, aligning assessment with the business goal of credit scoring. It also supports profit-based feature selection by measuring each feature’s contribution to model profit.

  • Motivation: Traditional credit-scoring metrics assess discrimination, whereas EMP incorporates expected losses and operational income.EMP is tailored to the business goal of credit scoring and facilitates computing model value.
  • Parameterization: EMP depends on benefit, classification cost, action cost, and the distributions of defaulters and non-defaulters, with parameters derived from loan economics.The framework uses LGD, EAD, loan amount, ROI, and an estimated uncertainty distribution for λ.
  • Measure: EMP integrates classification profit over possible costs and identifies the cutoff-dependent maximum profit.The optimal cutoff determines the fraction of applications that should be rejected to maximize profit.
  • Model profit: Model profit is computed by assigning defaulter or non-defaulter labels at a cutoff, applying the confusion matrix, and aggregating customer-level gains and losses.The cutoff depends on the number of test-set instances.
  • Feature selection: Profit-based feature selection ranks features by mean decrease in profit across random-forest trees.The procedure compares average profit for trees containing a feature with average profit for trees where that feature is absent.

4. Experimental Design

The study combines anonymized telecommunications and banking data to construct credit-scoring features for card applicants. Call-detail records are aggregated into timeframe-specific networks, from which calling, link-based, and influence-propagation features are extracted alongside traditional bank features.

  • Data sources: The dataset combines five months of telecommunications CDR data with anonymized bank data covering demographics, debit activity, and credit-card activity.The telco data contains almost 90 million unique cell phone numbers, while the bank data includes over two million customers.
  • Target and bank features: Credit-card activity supplies payment arrears, credit limits, remaining credit, and repayment outcomes for predicting creditworthiness and computing EMP.Customers are followed for up to twelve months after receiving the card.
  • Network construction: Three call networks are built by aggregating three months of CDRs before each card-acquisition month and retaining links for calls lasting at least five seconds.Subjects are considered separately across three timeframes.
  • Network composition: Network nodes include card-receiving subjects, other bank customers, and telco-only customers, while bank customers with arrears are marked as delinquent.The network distinguishes subjects, delinquent bank customers, non-delinquent bank customers, and telco customers who are not bank customers.
  • Network features: The analysis extracts calling-behavior, link-based, PageRank, and SPA exposure features, using delinquency labels with one, two, or three late payments.Edges are weighted by call counts and incoming, outgoing, and undirected relationships are considered.

5. Results

The call networks exhibit statistically significant relational structure around default, providing a basis for using social-network features to predict credit default. However, the reported network pattern includes both evidence of defaulter homophily and lower-than-expected cross-label connectivity.

  • Network dependency: A one-tailed proportion test finds evidence of homophily among defaulters with p-value less than 0.0001.The test uses a normal approximation.
  • Network dependency: Defaulters show dyadicity of 0.8689 and heterophilicity of 0.8137, indicating that the analyzed networks are not dyadic.Dyadicity measures same-label connectedness, while heterophilicity measures different-label connectedness relative to random-network expectations.
  • Implication: The observed relational dependencies provide a foundation for applying social-network analytic techniques to predict default in call networks.The results are organized around empirical tests of network relational dependency before model-performance analysis.

F G H

Random forests outperform the other classifiers, and network-related features contribute strongly to both statistical and economic credit-scoring performance. Calling behavior is especially predictive, while profit rankings depend on the feature-group combination and economic assumptions.

  • Statistical performance: Random forests produce the best-performing models, while logistic regression performs worst and does not improve with network-related features.The results suggest that generalized linear models do not capture the relevant nonlinear behavior.
  • Feature importance: Calling-behavior features rank highest in model H’s mean-decrease-in-accuracy importance analysis, followed by PageRank and SPA features.One link-based feature also appears among the 20 most important variables.
  • Economic assumptions: Economic results are robust to LGD variation, but expected maximum profit decreases as ROI increases; the analysis sets LGD to 0.8 and ROI to 0.05.The ROI value is selected from the elbow of the sensitivity analysis.
  • Economic performance: Expected maximum profit rankings are consistent with AUC rankings, with models A, F, G, and H best and model C worst.The profit-maximizing rejection fractions vary across models even when expected maximum profit rankings agree.
  • Economic performance: Model B, using calling-behavior variables alone, has the highest profit among single-source models, although model A follows marginally.The authors suggest immediate network information may reveal socioeconomic standing when borrower history is unavailable.
  • Profit-based importance: Among combined models, F and H produce the best profits, while calling-behavior features comprise more than half of the profit-important features.Sociodemographic consumption features comprise roughly a quarter and align with model F’s larger profit.
  • Interpretation: Network-only features can discriminate customers not captured by common socioeconomic and demographic variables, and these customers bring substantial profit.The reported association concerns customer discrimination and profitability rather than a causal mechanism.

6. Discussion

Call-detail records add predictive value to credit scoring, while calling behavior can outperform traditional data economically and statistically. Network features also provide evidence about how default influence propagates, although that propagation requires further analysis.

  • Models using all features performed best statistically, whereas model B achieved the highest profit and was slightly better than traditional model A.
  • AUC increased by 0.023 points in the best model compared with sociodemographic model A.
  • Models combining feature groups produced lower profit but higher EMP because their higher EMP fraction excluded more defaulters.
  • Call data had predictive power at least as good as traditional data for these borrowers, as shown by model B’s high performance.
  • Calling behavior features were more predictive than traditional features and enabled borrowers with limited bank information to obtain credit through call-record sharing.
  • Homophily tests found fewer connections among defaulters and between defaulters and non-defaulters, possibly because few defaulters were present overall.
  • Count Low Exposure features indicate that lacking a high-risk neighbor predicts non-default, while PageRank exposure features capture influence from delinquent customers.
  • Personalized PageRank was more effective than Spreading Activation for propagating default influence, but more analysis is needed to distinguish their effects.

7. Impact of Research

The findings have implications for regulatory risk measurement, financial inclusion, privacy, and ethical data use. Call data can broaden credit access, but sharing and social-network-based decisions remain constrained by regulation and fairness concerns.

  • Regulatory implications: The study illustrates default behavior propagating through a call network and proposes further research to quantify retail asset and default correlations.
  • Regulatory implications: Such research could support more empirically grounded regulatory asset-correlation values and better protection of the financial system.
  • Financial inclusion: Call data may facilitate credit access for people lacking banking history, particularly where historical financial data is often nonexistent.
  • Financial inclusion: Untraditional data features were good predictors of credit behavior, including in models B, G, and H.
  • Privacy and data sharing: Data sharing between telcos and financial institutions is constrained by the absence of worldwide standards and differing regulatory frameworks.
  • Privacy and data sharing: In the European Union, CDR-based financial scoring is treated as a value-added service requiring explicit user authorization unless the data is anonymized.
  • Ethical use: Using social-network data to restrict funding can constitute unfair discrimination, although call data may contribute to inclusion when conventional behavior information is unavailable.
  • Ethical use: The authors propose using call data in strictly positive terms to facilitate inclusion for people with insufficient information.

8. Conclusion

The study finds that mobile-phone calling behavior and social-network analytics can increase the statistical and economic value of credit-scoring models. Calling-behavior features perform best across AUC and profit, while the results support using mobile-phone data in positive terms to facilitate financial inclusion.

  • Incorporating telco data has the potential to increase the Value of credit-scoring models from both statistical and profit perspectives.The analysis uses social-network analytics and call networks built from phone-call logs for creditworthiness prediction.
  • Calling-behavior features perform best in both AUC and profit, and dominate other features in predictive importance.This pattern holds across models evaluated for statistical and economic performance.
  • Using phone behavior as the sole data source could support loan decisions for applicants lacking sufficient traditional information.The authors propose that such data be used in strict positive terms to facilitate financial inclusion.
  • The evidence is limited because the scorecards concern credit-card applications, and generalization to microloans or mortgages remains unclear.The study’s data come from a single-country agreement between a telco and a bank, and credit-bureau variables such as FICO scores were unavailable.
  • The resulting models show positive effects on financial inclusion and model profit.The paper frames this added value as the fifth V of Big Data: Value.
Loading 2002.09931v1…