Source-linked AI summary

A Deep Causal Inference Approach to Measuring the Effects of Forming Group Loans in Online Non-profit Microfinance Platform

Thai T. Pham, Yuanyuan Shen

arXiv:1706.02795v1stat.MLcs.IRcs.LGq-fin.GN

TL;DR

The paper asks whether forming group loans accelerates funding on Kiva, a philanthropic online marketplace where conventional group-lending arguments may not directly apply. It combines causal inference with deep learning to incorporate unstructured loan-description text, finding that group loans fund about 3.3 days faster on average. The authors conclude that field partners should generally encourage borrowers to form groups, while noting that effects may differ across borrower and loan categories.

  • Problem

    The paper asks whether forming group loans affects funding time on Kiva’s philanthropic online crowdfunding platform, where lender behavior differs from traditional microfinance and online lending.

  • Method

    The paper combines causal inference with deep learning to estimate treatment-effect components from high-dimensional loan-description text and Kiva project data.

  • Results

    3.3 days fewer is the reported average funding time for group loans than for individual loans on Kiva.

  • Takeaways & Limitations

    Field partners are advised to pool borrowers into groups, and group lending is described as enabling lenders to diversify their loan portfolios.

  • Takeaways & Limitations

    Future work should estimate average effects for borrower subgroups such as agriculture, food, and retail loans because treatment effects may differ across subgroups.

Abstract

from arXiv · show

Kiva is an online non-profit crowdsouring microfinance platform that raises funds for the poor in the third world. The borrowers on Kiva are small business owners and individuals in urgent need of money. To raise funds as fast as possible, they have the option to form groups and post loan requests in the name of their groups. While it is generally believed that group loans pose less risk for investors than individual loans do, we study whether this is the case in a philanthropic online marketplace. In particular, we measure the effect of group loans on funding time while controlling for the loan sizes and other factors. Because loan descriptions (in the form of texts) play an important role in lenders' decision process on Kiva, we make use of this information through deep learning in natural language processing. In this aspect, this is the first paper that uses one of the most advanced deep learning techniques to deal with unstructured data in a way that can take advantage of its superior prediction power to answer causal questions. We find that on average, forming group loans speeds up the funding time by about 3.3 days.

1 Introduction

The paper studies whether forming group loans on Kiva accelerates funding in a philanthropic online marketplace. It combines deep learning for loan-description text with causal inference and estimates that group loans fund about 3.3 days faster on average.

  • Kiva and the funding problem: Kiva connects philanthropic lenders worldwide with borrowers through field partners that screen applicants, post requests, disburse loans, and collect repayments.Field partners often pre-disburse loans, making rapid fundraising important.
  • Group loans: Group lending assigns members responsibility for one another’s loans, potentially reducing lenders’ risk exposure through shared liability and borrower pooling.Traditional microfinance institutions encourage groups because borrowers who know one another may be safer and more likely to be approved.
  • Research question: Kiva’s nonprofit lenders can diversify across many loans, so conventional repayment-based arguments for group lending may not apply directly in this setting.Lenders can choose their investment amounts, subject to a $25 minimum.
  • Research question: The paper measures the average treatment effect of forming group loans on the time required for projects to become funded.The central question is whether field partners should organize borrowers into groups to accelerate funding.
  • Main result: 3.3 days faster funding is the estimated average effect of group loans, with cutting-edge methods producing roughly −3.3 estimates and a standard deviation of 0.167.The result is reported as a significant negative treatment effect on funding time.
  • Methodological contribution: Deep causal inference combines deep learning and causal inference to use unstructured loan-description text when estimating causal effects.The paper presents this as its first use of an advanced deep learning technique for unstructured data in causal questions.

2 Literature Review

Prior research explains group lending through repayment, monitoring, contract, and adverse-selection mechanisms, while online-platform research examines funding behavior in different marketplace settings. This paper instead studies funding time for group versus individual loans on a philanthropic online crowdfunding platform and uses deep learning to incorporate textual data into causal analysis.

  • Online crowdfunding microfinance: Online crowdfunding platforms differ in what funders receive, ranging from tangible rewards on Kickstarter and Indiegogo to principal and interest on Prosper and Lending Club.Prior work in these settings studies project success, failure, and funding dynamics.
  • Online crowdfunding microfinance: Kiva’s nonprofit-driven lenders respond to borrower narratives, making the platform distinct from conventional online lending marketplaces.Prior Kiva research found that narratives framing businesses as helping others receive more positive responses than opportunity-focused narratives.
  • Group-lending literature: Traditional microfinance theory links group loans to reduced moral hazard, improved commitment, and mitigation of adverse selection.These mechanisms operate through joint liability, monitoring, contract structure, and borrower self-selection.
  • Group-lending literature: Empirical studies in Thailand, Guatemala, and Burkina Faso examine repayment performance, peer monitoring, social ties, and group dynamics.These studies focus on traditional microfinance groups rather than online funding speed.
  • Paper’s contribution: This paper differs by estimating effects on time until projects are funded rather than borrower repayment, using online crowdfunding instead of traditional MFIs.Its outcome and institutional setting therefore extend the empirical group-lending literature.
  • Causal inference and deep learning: The paper combines causal estimators with deep learning to process high-dimensional, unstructured economic data such as text.It considers Double Selection, Doubly Robust, and Targeted Maximum Likelihood estimators, while emphasizing textual data from Kiva loan descriptions.

3 Data

The study analyzes Kiva loan records to compare group and individual borrowing, using borrower descriptions and other covariates to examine funding outcomes. It describes loan composition, sectors, gender, funding times, and loan amounts.

  • 995,911 Kiva loan entries from January 1, 2006 to May 10, 2016 form the cleaned dataset.
  • The outcome is funding time, while the treatment is whether borrowers form a group loan rather than request an individual loan.
  • Borrower loan descriptions are included alongside other information as covariates for the analysis.
  • Group loans average $1816, compared with $652 for individual loans.
  • Housing has the longest average funding time at 11.55 days, while arts has the shortest at 1.57 days.
  • Wholesale loans average $1,228, the largest amount among sectors, while personal-use loans average $541.

4 Causal Inference Setting

The paper frames group formation as a binary treatment and funding time as the outcome within the potential-outcome framework. Identification relies on standard assumptions concerning observed covariates, treatment overlap, and interference.

  • The potential-outcome framework represents covariates as X, group formation as binary W, and funding time as outcome Y.
  • The target estimand is the average treatment effect τ = E[Y(1) − Y(0)].
  • SUTVA assumes that one project's funding outcome is unaffected by other projects' treatment decisions and that treatment has constant value.
  • Unconfoundedness requires potential outcomes to be independent of treatment conditional on observed covariates.
  • The exogeneity assumption rules out unobserved covariates jointly affecting group formation and funding time, but the authors note it is almost impossible to validate.
  • Overlap requires 0 < P(W = 1|X) < 1, ensuring observations exist in both treatment groups for every covariate value.

5 Preliminary Analysis

The preliminary analysis compares funding times for group and individual loans while examining loan amount, sector, and treatment-group differences. These comparisons motivate more advanced methods for estimating the treatment effect.

  • Overall comparison: 0.17 days is the naive estimated ATE, with treated loans averaging 7.25 days and control loans 7.08 days.The estimate has a standard deviation of 0.027.
  • Loan amount adjustment: 0.15 versus 0.31 average days per $25 indicates faster funding for group loans after scaling by loan amount.The group-loan ratio suggests less than 4 hours per $25, compared with more than 7 hours for individual loans.
  • Loan amount adjustment: The linear regression implies negative treatment effects above $390 in loan amount but positive effects below $390.The low R2 indicates that the linear model does not fit the data well, motivating methods that model nonlinearities.
  • Sector heterogeneity: Group effects vary across sectors from negative to neutral to positive among funded loan requests.Figure 7 reports average funding time by sector and treatment group.
  • Sector composition: Food, agriculture, and retail are the three largest loan categories for both group and individual loans, reducing concern that sector composition explains funding-time differences.Entertainment and wholesale are the least frequent categories for both treatment groups.

6 Methodology

The methodology combines causal-inference estimators with machine-learning models to incorporate high-dimensional and textual loan covariates. It compares a regularized linear baseline with DSE, DRE, and TMLE, including deep-learning estimates of nuisance components.

  • Motivation: High-dimensional structured and unstructured data create challenges for causal inference, motivating machine-learning methods for treatment-effect estimation.The Kiva loan descriptions are difficult for traditional econometric models to use without discarding information.
  • Preprocessing: GloVe or Word2Vec transforms loan descriptions into word-level vectors that can be combined into high-dimensional loan covariates.These processed vectors are used in both baseline and advanced models.
  • Baseline: The baseline uses elastic-net linear regressions to estimate treated and control outcome relations from 17 non-text covariates.Predicted treated and control outcomes are then used to estimate the treatment effect.
  • Deep causal inference: Deep-learning models estimate the infinite-dimensional components of DRE and TMLE using either MLPs with processed covariates or recurrent models with word embeddings.The text representations are incorporated with the other covariates before causal estimation.
  • Advanced estimators: DSE uses Lasso-selected covariates explaining outcomes or treatment, then estimates treated and control outcomes separately with OLS.The method is described as computationally inexpensive and reasonably effective in many cases.
  • Advanced estimators: DRE and TMLE estimate outcome and propensity-score components, with DRE remaining consistent if either model is correctly specified.TMLE updates the initial component estimates while retaining double robustness.

7 Deep Learning Techniques

The paper uses GloVe embeddings with MLP and Deep LSTM architectures to represent loan descriptions for causal-model components. MLP averages word vectors, whereas Deep LSTM preserves sequential relations in the text.

  • Overview: Deep learning is used because numerical representations of Kiva text data are high-dimensional and suited to these models.The paper describes deep learning as especially effective for prediction when data are abundant.
  • Text representation: GloVe creates word embeddings intended to capture semantic and syntactic relations.The paper uses pretrained GloVe vectors rather than retraining them on Kiva data.
  • MLP: MLP preprocessing averages all word embeddings in a loan description into one representative loan vector.The resulting vector feeds the baseline or MLP model.
  • MLP: The MLP combines a 100-dimensional loan-vector input with a separate layer of 17 other covariates.The architecture uses hidden layers, regularization, and an output layer for prediction.
  • Deep LSTM: Deep LSTM processes a variable-length sequence of d-dimensional word vectors and preserves sequential relations that MLP averaging neglects.The recurrent output is passed into additional fully connected layers with the other covariates.
  • Deep LSTM: LSTM cells address vanishing or exploding gradients by controlling which terms participate in multiplications.For outcome models, the architecture uses ReLU, additional layers, and attention over recurrent outputs.

8 Results

The paper compares baseline and deep-learning models for estimating propensity scores, outcomes, and average treatment effects using loan text and covariates. Deep-learning approaches generally perform better for component estimation, while treatment-effect estimators consistently indicate that group loans reduce funding time.

  • Covariate Relatedness Check: 83 of 100 text-embedding covariates are significant in the treated outcome model, compared with 92 in the control model.These results indicate that loan text correlates with the outcome variable.
  • Covariate Relatedness Check: All but four covariates are significant in the treatment model, while all but seven are significant in the outcome model.The authors interpret these models as dense and text-dependent.
  • Estimation Model Comparison: Deep LSTM performs particularly well for propensity-score estimation, while MLP and Deep LSTM outperform RLR and RF for outcome estimation.The outcome-model advantage is smaller, possibly because extreme test-set outliers affect every model.
  • Estimation Model Comparison: Including text data improves every model estimation and substantially boosts the propensity-score F1 scores of RLR and RF.The text data is especially important for estimating propensity scores.
  • Comparison of Average Treatment Effect Estimators: Naive and baseline estimators can imply no effect or increased funding time, unlike advanced deep-learning approaches.The baseline model produces a statistically significant positive ATE, opposite to the advanced-method results.
  • Comparison of Average Treatment Effect Estimators: DREs with deep learning give almost identical ATE estimates, while TMLEs with deep learning differ in magnitude but all indicate a significantly negative effect.DRE and TMLE are treated as reliable because of double robustness, while observations with estimated propensity scores outside [0.01, 0.99] are excluded.

9 Conclusion

The conclusion emphasizes deep learning’s role in incorporating unstructured text into causal estimation and reports that group loans receive funding faster on average. It also identifies subgroup-specific treatment effects as a direction for future work.

  • Contribution: Deep learning models outperform other approaches in estimating propensity-score and outcome-model components for causal estimators.The paper combines deep learning with causal inference to estimate treatment effects using unstructured text data.
  • Main Result: 3.3 days fewer is the average funding time for grouped loans than for individual loans on Kiva.The paper describes this as a significant treatment effect and advises field partners to pool borrowers into groups.
  • Future Work: Future work will estimate average effects separately for borrower subgroups such as agriculture, food, and retail loans.The authors expect treatment effects to differ across these subgroups.

Appendix A Raw Data Description

The appendix describes a Kiva loan dataset retrieved between January 1, 2006 and May 10, 2016, stored across 2,240 JSON files. It lists the raw attributes retained for each loan record.

  • Dataset: 2,240 JSON files contain Kiva loan data retrieved from January 1, 2006 to May 10, 2016.The project uses only loan data from these files.
  • Raw Attributes: Raw records include borrower, loan, funding, repayment, location, partner, sector, status, description, and timing attributes.The listed fields include funded_amount, loan_amount, posted_date, funded_date, borrowers, description, location, partner_id, sector, and status.

A.3 Sample Raw Data

The appendix presents one raw Kiva record for Mahesh, an Indian education borrower seeking funds for postgraduate course fees. The example combines narrative, financial, timing, location, and loan-term fields.

  • Borrower and Use: The sample borrower is Mahesh, whose activity is higher education and whose stated use is PGDBM course fees.The record includes an English narrative about studying and improving quality of life.
  • Funding Details: $1,150 is the loan amount, with 46 lenders and a funded date of 2015-03-24.The raw record also reports funded_amount 1150, loan id 853701, and loan location information.
  • Posting and Status: The loan was posted on 2015-03-18, listed in India’s Education sector, and marked funded.The record identifies the borrower as Mahesh and includes a planned expiration date of 2015-04-17.

Appendix B Data Manipulation

The study constructs treatment, outcome, text, and covariate data by filtering Kiva loan records and encoding borrower, risk, sector, and language information. It then uses normalized covariates and deep learning models with pretrained word representations to estimate the models.

  • Data Manipulation: The preprocessing removes unusable, post-decision, redundant, or non-text fields and retains only loans with actual English descriptions.Excluded fields include covariates with missing or identical values, hard-to-process or seemingly unimportant fields, and variables occurring after funding decisions.
  • Data Manipulation: Group-loan treatment W equals 1 for loans with multiple borrowers and 0 for loans with one borrower; examples without borrowers are discarded.The treatment is created by counting borrowers.
  • Data Manipulation: Funding time Y is the difference between funded_date and posted_date in days, while never-funded loans are excluded because their outcome is infinity.Loans already funded before posting are also discarded.
  • Covariate Construction: Gender is assigned by majority borrower gender, risker records whether the partner or lender bears default risk, and 14 sector dummies replace the 15-category sector variable.The original borrower-gender, risker, and sector variables are transformed into analysis-ready covariates.
  • Data Manipulation: 995,911 loan entries remain after preprocessing, with description_texts and 17 covariates including sector dummies, loan_amount, risker, and gender.loan_amount is the only nonbinary covariate and is normalized to zero mean and unit variance.
  • Modeling: The models use fixed 100-dimensional Wikipedia-trained GloVe vectors, averaging word vectors for the Multilayer Perceptron but retaining word-level embeddings for Deep LSTM.Both L2 loss and dropout are used for regularization, and training is performed in TensorFlow with GPUs.
Loading 1706.02795v1…