Source-linked AI summary
To Explain or to Predict?
Galit Shmueli
TL;DR
The paper addresses the common conflation of causal explanation with empirical prediction in statistical modeling. It compares the two goals across the modeling process and concludes that they require different choices and should be evaluated separately. The distinction has implications for scientific theory development and practice.
Problem
Statistical fields often assume that models with high explanatory power also have high predictive power, while the broader statistical literature lacks a systematic treatment of their distinction.
Method
The article examines statistical modeling from goal definition through model use and reporting, comparing methods, criteria, data, and information for explanatory versus predictive aims.
Results
Explanatory and predictive goals produce different modeling choices and final models, with predictive modeling providing scientifically useful information beyond causal explanation.
Takeaways & Limitations
Researchers should specify the study goal a priori and report both explanatory and predictive qualities rather than treating predictive validity as evidence that a model is truer.
Takeaways & Limitations
The article notes that some unjustified optimism remains when repeated statistical-learning procedures and variable transformations are not fully accounted for in AIC-based model selection.
Abstract
from arXiv · showhide
Statistical modeling is a powerful tool for developing and testing theories by way of causal explanation, prediction, and description. In many disciplines there is near-exclusive use of statistical modeling for causal explanation and the assumption that models with high explanatory power are inherently of high predictive power. Conflation between explanation and prediction is common, yet the distinction must be understood for progressing scientific knowledge. While this distinction has been recognized in the philosophy of science, the statistical literature lacks a thorough discussion of the many differences that arise in the process of modeling for an explanatory versus a predictive goal. The purpose of this article is to clarify the distinction between explanatory and predictive modeling, to discuss its sources, and to reveal the practical implications of the distinction to each step in the modeling process.
1. INTRODUCTION
The article distinguishes explanatory modeling, which targets causal explanation, from predictive modeling, which targets empirical prediction. It argues that conflating them affects statistical practice, scientific theory building, and the use of models.
- 1. INTRODUCTION: The distinction matters because explanatory power does not inherently imply predictive power.Both forms are useful for generating and testing theories, but they play different roles.
- 1. INTRODUCTION: Explanatory and predictive modeling pursue distinct scientific goals: causal explanation versus empirical prediction.The paper also distinguishes descriptive modeling, which parsimoniously captures data structure.
- 1. INTRODUCTION: The article examines how explanatory and predictive goals change methods, criteria, data, and information throughout the modeling process.Its scope extends from goal definition and study design through model use and reporting.
- 1.1 Explanatory Modeling: Explanatory modeling applies statistical models to test causal hypotheses about theoretical constructs, often translating constructs into measurable variables.The illustrated information-systems workflow begins with theory, causal hypotheses, diagrams, and construct operationalization.
- 1.4 The Scientific Value of Predictive Modeling: Predictive modeling can uncover complex patterns, suggest new causal mechanisms and measures, improve explanatory models, compare theories, and assess theory–practice distance.The paper presents predictive accuracy as scientifically informative even when prediction is not the study’s primary goal.
- 1.6 A Void in the Statistics Literature: The article addresses a literature gap by examining the explanatory–predictive distinction across the broader statistical modeling process, not only model selection.It focuses particularly on implications for theory development by nonstatistician scientists.
2. TWO MODELING PATHS
The paper treats statistical modeling as a sequence of decisions whose methods differ according to whether the goal is explanation or prediction. These differences produce distinct final explanatory and predictive models.
- 2. TWO MODELING PATHS: Modeling choices differ between prediction and explanation in methods, criteria, data, and information considered at each step.The process spans goal definition through model use and reporting.
- 2. TWO MODELING PATHS: The main study goal should be specified early because explanatory and predictive objectives optimize different criteria.Descriptive modeling can also be a study goal, but the article focuses on explanation and prediction.
- 2. TWO MODELING PATHS: Conceptual and practical differences across the modeling process lead to different final explanatory and predictive models.The distinction is therefore procedural, not merely a difference between two labels for the same fitted model.
2.1 Study Design and Data Collection
Study design and data collection must be tailored to the modeling goal. Explanation prioritizes inference about causal structure, whereas prediction prioritizes accurate performance on future observations and realistic contexts.
- Study Design and Data Collection: Predictive studies generally require more data than explanatory studies because estimating f from data, creating holdouts, and predicting individuals add uncertainty.Explanatory studies emphasize statistical power and model-specification testing, with diminishing inferential returns beyond sufficient precision.
- Study Design and Data Collection: Hierarchical sampling should allocate observations differently for estimation and prediction.For prediction, increasing group size n may help more; for estimation, increasing the number of groups J may help more.
- Study Design and Data Collection: A theoretically appropriate hierarchical f can have poorer predictive performance than a nonhierarchical f in hierarchical data.This illustrates that theoretical appropriateness and predictive performance need not coincide.
- Study Design and Data Collection: Experimental data are preferred for causal explanation, while observational data can be preferable for prediction when they better represent realistic uncontrolled conditions.Prediction emphasizes prospective representativeness, including realistic noise and measured responses.
- Study Design and Data Collection: Measurement instruments should represent underlying constructs adequately for explanation, but predictive work emphasizes measurement quality and relevance to the variable being predicted.The criteria differ because the modeling goals differ.
- Study Design and Data Collection: Factorial designs target causal factors with interpretable linear models, whereas response-surface designs target predictor combinations that optimize Y using nonlinear estimation.The contrast links experimental-design choices to explanatory versus predictive objectives.
2.2 Data Preparation
Missing-value handling and data partitioning illustrate how data preparation depends on whether the goal is explanation or prediction. Predictive preparation prioritizes performance on unseen or incomplete future cases.
- Data Preparation: Missing-data solutions depend on whether missing values occur in training data or in cases that must be predicted.Prediction cannot simply discard cases with missing inputs.
- Data Preparation: Missingness indicators may be unsatisfactory for explanatory modeling yet produce excellent predictions.Their predictive usefulness is strongest when missingness is informative about Y, such as missing financial data associated with fraudulent reporting.
- Data Preparation: For prediction with varying missing predictor sets, multiple reduced models can be estimated, each omitting a different subset of predictors.The model used for an observation depends on which predictor information is missing.
- Data Preparation: Holdout samples, cross-validation, and resampling help avoid overoptimistic predictive accuracy from evaluating on training data.These procedures assess performance on data not used to build the model.
- Data Preparation: Data partitioning trades some bias for reduced sampling variance and is therefore more useful for predictive than explanatory modeling.In explanatory work, partitioning reduces statistical power and is less common, though it can assess robustness or provide predictive validation.
2.3 Exploratory Data Analysis
Exploratory data analysis supports both explanatory and predictive modeling, but its focus differs: theory-guided causal relationships versus broad investigation of associations, measurement quality, and predictive inputs.
- EDA summarizes, visualizes, reduces, and prepares data for formal modeling in both explanatory and predictive contexts.
- Explanatory EDA is guided by theoretically specified causal relationships, whereas predictive EDA explores measurement quality and associations more broadly.
- Exploratory visualization is interactive and open-ended, while confirmatory visualization tests a specified hypothesis under more predetermined conditions.
- Predictive analysis may examine numerical summaries across many variables, whereas explanatory analysis focuses summaries on theoretically relevant relationships such as mediation.
- Predictive modeling may use PCA or other less interpretable compression methods to reduce sampling variance, while explanatory exploration is more restrictive.
2.4 Choice of Variables
Variable choice differs because explanatory modeling assigns variables roles within a causal theory, whereas predictive modeling prioritizes useful, high-quality, ex-ante available associations.
- Variable-selection criteria differ markedly between explanatory and predictive contexts.
- Explanatory variables are chosen according to their construct’s role in the theoretical causal structure and the quality of its operationalization.
- Explanatory frameworks distinguish roles such as antecedent, consequent, mediator, moderator, treatment, control, exposure, and confounder.
- Omitting true health status from a health-insurance model can create endogeneity through reverse causation or measurement error in insurance status.
- Predictive selection emphasizes association quality, data quality, and whether predictors are available before prediction, without requiring an exact causal role.
2.5 Choice of Methods
Explanatory modeling favors interpretable methods linked to a theoretical causal model, whereas predictive modeling permits broader statistical and algorithmic methods chosen for accurate future predictions.
- Explanatory models require interpretable statistical functions that can be linked to the underlying theoretical model.
- Predictive modeling includes interpretable and uninterpretable statistical models as well as data-mining algorithms because the data-generating function is often unknown.
- Algorithmic modeling is suitable for predictive and descriptive modeling but not for explanatory modeling.
- A contemporaneous oil-price model may explain airfare causally but cannot predict future airfare when oil price at prediction time is unknown.
- Prediction therefore may require lagged predictors or forecasts of unavailable inputs, whereas centered moving averages are unusable when future-window data are unavailable.
- Shrinkage methods reduce estimation variance for prediction by introducing bias, even though they are not intended for explanation.
2.6 Validation, Model Evaluation and Model Selection
Validation, evaluation, and selection serve different goals: explanatory modeling tests theoretical representation and inference, while predictive modeling emphasizes generalization to new data and calibrated predictive performance.
- Validation: Explanatory validation checks whether the model represents the theory and fits observed data, whereas predictive validation assesses generalization to new observations.
- Validation: Predictive validation compares training and holdout performance to detect overfitting, the principal threat to generalization.
- Validation: Multicollinearity threatens coefficient inference and attribution in explanatory models but does not affect predictive ability.
- Model evaluation: Explanatory performance uses relationship strength and significance measures such as R2 and F, whereas predictive performance uses out-of-sample metrics tailored to the prediction task and error costs.
- Model evaluation: R2 and F indicate association rather than causation, so predictive power cannot be inferred from explanatory power and the two should be assessed separately.
- Model evaluation: Predictive metrics should be computed on holdout data or through cross-validation because fitted-data performance tends to overestimate predictive accuracy.
- Model selection: Model selection retains theoretically meaningful explanatory variables when justified, while predictive selection can favor removing small-coefficient inputs and using predictive criteria.
- Model selection: Repeated tuning and selection can make predictive metrics overoptimistic because criteria such as AIC ignore information accumulated across prior fitting attempts.
2.7 Model Use and Reporting
Explanatory and predictive models serve different purposes and use fitted functions differently. Explanatory studies emphasize inference and causal conclusions, whereas predictive studies use fitted functions to generate predictions for new data and compare predictive power across models.
- Explanatory and predictive models differ in their data, estimated function, explanatory power, predictive power, and use.These differences arise throughout the modeling process rather than only in the final fitted model.
- Explanatory models use inference to derive statistical conclusions that are translated into scientific conclusions about causal hypotheses.Their focus includes theory, causality, bias, and retrospective analysis.
- Predictive models use the fitted function to generate predictions for new data.The difficulty depends on the function's complexity and the type of prediction, such as a complete predictive distribution.
- Predictive scientific studies emphasize observable data, association, bias–variance considerations, and prospective analysis.Their conclusions can address hypothesis generation, practical relevance, and predictability level.
3. TWO EXAMPLES
The examples show that predictive and explanatory studies make different choices about data, variables, methods, evaluation, and reporting. Predictive models can improve accuracy and generate hypotheses, but they do not by themselves provide direct causal explanations.
- Examples: The examples convert predictive and explanatory studies into the other mode to illustrate how their modeling choices differ.The paper considers both a predictive goal converted to an explanatory study and an explanatory study converted to a predictive context.
- Netflix Prize: The Netflix contest targeted accurate prediction of unseen user–movie ratings, using a very large dataset and a $1,000,000 grand prize.The task required improving predictive accuracy over Netflix’s recommendation engine by at least 10%.
- Netflix Prize: Netflix’s winning approach combined nearest-neighbor, regression, and shrinkage methods, with blended models producing more accurate predictions.The passage describes collaboration between competing teams as another source of improved accuracy.
- Netflix Prize: The Netflix example shows that predictive modeling can have scientific value by revealing predictability, rating-scale usefulness, and informative missingness.It also sets the stage for explanatory research.
- Netflix Prize: A predictive study of movie preferences would require causal hypotheses, theoretical constructs, operationalization, and covariates capturing movie and user characteristics.The predictive dataset alone would not determine the explanatory study’s causal structure.
- Online Auction Research: Explanatory auction research uses theory-based constructs and regression models to study factors affecting final auction prices.One example used 461 eBay coin auctions and modeled bidder valuation and other constructs.
- Online Auction Research: Explanatory auction modeling retains variables and reports inference to match theoretical constructs with causal hypotheses.Insignificant covariates may be retained for matching the fitted function with the theoretical function, while conclusions are expressed causally.
- Online Auction Research: Predictive auction studies choose variables according to forecast timing and evaluate models with out-of-sample metrics such as MAPE and RMSE against benchmarks.Variables unavailable at the prediction time cannot be used, whereas information accumulated during an auction can be useful for later forecasts.
4. IMPLICATIONS, CONCLUSIONS AND SUGGESTIONS
The paper argues that explanatory power and predictive accuracy are distinct dimensions, so scientific modeling must specify its goal and report both qualities. Distinguishing them can improve theory development, statistical practice, and the connection between research and practice.
- Scientific research and practice: Neglecting predictive modeling alongside explanatory modeling can weaken theory testing, hinder discovery of causal mechanisms, and widen the gap between research and practice.The paper links these consequences to fields that rely nearly exclusively on causal explanation and omit predictive testing.
- Scientific research and practice: High explanatory power does not establish predictive power, and treating in-sample measures such as R2 or hazard ratios as predictive evidence can produce incorrect conclusions.Examples are described in ecology, economics, epidemiology, and information systems.
- Two dimensions: Explanatory power and predictive accuracy should be treated as two dimensions because models can possess different levels of each.The paper recommends reporting both qualities even when only one is the study’s primary goal.
- Two dimensions: Researchers should specify the modeling goal a priori and align model development and evaluation with the criterion they intend to optimize.The proposed two-dimensional framework also supports comparing models by both explanatory and predictive capabilities.
- Implications for model choice: The distinction makes parsimony task-dependent: a model overly complicated for explanation may be “sophisticatedly simple” for prediction when added complexity improves predictive accuracy.Predictive models may use complexity to capture small nuances in new data, whereas explanatory models may favor theoretical simplicity.
- Closing recommendations: The paper proposes integrating distinctions among explanatory, predictive, and descriptive modeling into statistics and research-methods education while promoting careful use of predictive terminology.The recommendations include accessible teaching materials and tools for implementing both explanatory and predictive modeling.
- Closing recommendations: Awareness of the different scientific functions of explanatory and predictive modeling is presented as essential for progressing scientific knowledge.The conclusion frames clearer distinctions as important for both scientific usage and statistical modeling practice.
- Implications for model choice: A less true, parsimonious model can have higher predictive validity than a truer but less parsimonious model.The appendix illustrates that an underspecified model can have greater bias but lower variance and, in some settings, lower expected prediction error.