Source-linked AI summary
Deep Learning in Asset Pricing
Luyang Chen, Markus Pelger, Jason Zhu
TL;DR
The paper addresses how to price individual stock returns using extensive conditioning information while retaining flexible, time-varying structure. It uses no-arbitrage moments, adversarial test-asset construction, and LSTM-extracted macroeconomic states. The resulting model outperforms benchmark approaches out of sample on Sharpe ratio, explained variation, and pricing errors.
Problem
The paper seeks to explain differences in average returns across individual assets using a stochastic discount factor that can incorporate extensive conditioning information.
Method
The model combines a no-arbitrage estimation criterion, adversarial selection of informative moment conditions, and recurrent extraction of hidden macroeconomic states.
Results
The asset-pricing model outperforms all benchmark approaches out of sample in Sharpe ratio, explained variation, and pricing errors.
Takeaways & Limitations
The framework identifies key factors that drive asset prices while incorporating flexible functional forms and time variation.
Takeaways & Limitations
The paper reports that macroeconomic data remain a limitation: even the most flexible model cannot fully compensate for this problem.
Abstract
from arXiv · showhide
We use deep neural networks to estimate an asset pricing model for individual stock returns that takes advantage of the vast amount of conditioning information, while keeping a fully flexible form and accounting for time-variation. The key innovations are to use the fundamental no-arbitrage condition as criterion function, to construct the most informative test assets with an adversarial approach and to extract the states of the economy from many macroeconomic time series. Our asset pricing model outperforms out-of-sample all benchmark approaches in terms of Sharpe ratio, explained variation and pricing errors and identifies the key factors that drive asset prices.
Related Literature
The paper places its deep-learning asset-pricing model within work using machine learning, no-arbitrage restrictions, flexible factor structures, and dynamic macroeconomic states. Its contributions combine an economic criterion, adversarial moment selection, and recurrent extraction of macroeconomic information.
- The paper contributes to machine-learning methods for asset pricing by imposing the no-arbitrage constraint during estimation.
- The paper reports better asset-pricing and explained-variation results than a simple prediction approach and identifies key factors driving asset prices.
- Its framework extends flexible neural-network asset-pricing approaches by identifying the stochastic discount factor rather than only predicting returns.
- An adversarial network selects moment conditions to handle the infinite number of moments implied by the no-arbitrage condition.
- Recurrent neural networks with Long-Short-Term-Memory units extract hidden states from macroeconomic time series to capture dynamic information.
B. Generative Adversarial Methods of Moments
The paper formulates asset-pricing estimation as a minimax method-of-moments problem. An adversary searches for portfolios and states with large pricing errors, while the model is adjusted to price them robustly.
- The SDF weights are estimated from conditional moment restrictions implied by no-arbitrage, with conditioning functions generating test assets.
- The adversarial approach chooses conditioning functions that produce the largest pricing errors against the candidate asset-pricing model.
- The minimax process alternates between finding mispriced portfolios or times and correcting the model until large pricing errors are no longer found.
- The framework addresses an infinite number of candidate moments and an infinite-dimensional parameter set, unlike conventional finite-dimensional GMM.
- D · N instrumented stocks provide a larger test-asset set than D portfolios, and this substantially accelerates convergence, empirically converging after three iterations.
A. Loss Function and Model Architecture
The model minimizes weighted pricing-error moments with neural networks that estimate SDF weights, adversarial instruments, and macroeconomic states. A feedforward benchmark instead predicts conditional mean returns.
- Loss Function: The empirical loss minimizes weighted sample moments, assigning greater weight to more precisely estimated moments and down-weighting short panels.
- Model Architecture: The SDF network uses a feedforward network to estimate portfolio weights that minimize pricing-error loss for a given conditioning function and information set.
- Model Architecture: A conditional network acts as an adversary, identifying assets and portfolio strategies that are hardest for the SDF network to explain.
- Model Architecture: An LSTM-based recurrent network summarizes macroeconomic information into state variables used with firm characteristics in the model.
- Alternative Model: The FFN benchmark estimates conditional mean returns by minimizing average squared prediction errors rather than using an adversarial network or LSTM.
C. Recurrent Neural Network (RNN) with LSTM
The paper uses an LSTM to extract low-dimensional, nonlinear macroeconomic state processes from current and past observations, capturing both cross-sectional dependence and long-term dynamics. These states condition the asset-pricing model without look-ahead bias.
- Motivation: Using only the latest differenced macroeconomic observation loses information and cannot identify cyclical dynamics such as business cycles.The paper motivates lagged inputs because the latest observation is insufficient to distinguish economic regimes.
- Implementation: The extracted state h_t depends only on current and past macroeconomic increments, avoiding look-ahead bias.The inputs are first transformed to stationary series before dynamic state extraction.
- Alternatives: PCA extracts factors describing innovation correlation but cannot identify the current economic state when that state depends on dynamics.The paper contrasts static dimension reduction with dynamic state extraction.
- Architecture: LSTMs are chosen because conventional RNNs can suffer exploding or vanishing gradients when relevant events lie far in the past.LSTM cells are designed for unknown and potentially long lags, making them suitable for business-cycle detection.
- LSTM states: LSTM states summarize high-dimensional macroeconomic series while retaining a general functional form and long-term dependencies.The authors describe the LSTM as combining factor-model-like cross-sectional aggregation with state-space-like dynamics.
- Evidence: In simulation, the LSTM successfully extracts a business-cycle pattern related to deviations of a local mean from a long-term mean.The empirical states also appear related to short- versus long-term averages of macroeconomic increments.
A. Data
The dataset combines long monthly stock-return histories, firm characteristics, and a broad panel of macroeconomic predictors. The authors standardize characteristics and transform macroeconomic series to obtain stationary inputs for estimation.
- Inputs: The data include monthly CRSP equity returns, 46 firm-specific characteristics, and 178 macroeconomic time series.The macroeconomic panel combines FRED-MD predictors, characteristic-median series, and Welch–Goyal predictors.
- Sample: The sample covers January 1967 to December 2016, with 20 years for training, 5 years for validation, and 25 years out of sample.The usable sample contains around 10,000 stocks after requiring complete firm-characteristic information.
- Inputs: Firm characteristics span past returns, investment, profitability, intangibles, value, and trading frictions.The characteristics are sourced from the Kenneth French Data Library or prior research datasets.
- Sample: The stock sample excludes observations missing any required firm-characteristic information.The authors note that imputing missing characteristics would introduce model-based assumptions and additional error.
- Preprocessing: Each characteristic is ranked cross-sectionally and converted into quantiles to normalize variables with different scales.The paper follows standard transformations used in related asset-pricing studies.
- Preprocessing: Macroeconomic variables and characteristic-median series are transformed using standard procedures to obtain stationary time series.The transformations follow McCracken and Ng (2016) for the macroeconomic data and are defined analogously for added series.
B. An Illustrative Example of GAN
The illustrative GAN example shows that asset-pricing performance depends jointly on SDF information and the conditioning variables used to construct test assets. Adding investment information to the test assets materially improves pricing across sorted portfolios.
- Model setup: GAN jointly learns SDF weights and a nontrivial conditioning function that constructs informative test assets.The example compares unconditional models with adversarial models using size, value, and investment information.
- Performance: Adding investment information to the test assets improves the model relative to GAN (SVI-SV), which omits investment from the conditioning function.The result shows that the characteristics included in test assets matter for SDF recovery, not only those in SDF weights.
- Performance: GAN (SVI-SVI) is roughly twice as good as UNC (SVI), depending on the performance metric.The comparison uses out-of-sample Sharpe ratio, explained variation, and cross-sectional R2.
- Conditioning function: The learned conditioning function emphasizes extreme portfolios involving small value stocks and large growth stocks, with investment further distinguishing conservative and aggressive firms.The paper interprets the learned structure as resembling Fama–French-type test assets.
- Portfolio pricing: GAN captures mean returns well across all sorted-portfolio quantiles, whereas UNC (SVI) fails particularly for small value stocks.The comparison uses model-implied versus average excess returns for 25 and 35 portfolios.
- Implication: The example demonstrates that estimating an asset-pricing model cannot be separated from choosing informative test assets.The conditioning function acts as a data-driven mechanism for identifying which assets matter for pricing.
C. Cross Section of Individual Stock Returns
In the individual-stock analysis, the GAN model outperforms benchmark approaches out of sample across Sharpe ratio, explained variation, and cross-sectional pricing performance. Hidden macroeconomic states and adversarially selected test assets are both important for this result.
- Model flexibility: The GAN’s nonlinear and interaction structure produces a 50% increase relative to the regularized linear model.The authors still report impressive performance for an appropriately designed linear model.
- Main results: GAN explains 8% of individual-stock-return variation, twice as much as the other models.This result concerns explained variation in the cross section of individual stocks.
- Main results: GAN obtains a cross-sectional R2 of 23%, substantially higher than the other models.The models are evaluated using out-of-sample Sharpe ratio, explained variation, and cross-sectional R2.
- Macroeconomic states: Including all 178 macroeconomic variables directly causes out-of-sample Sharpe-ratio performance to collapse across the models.Conditioning only on the latest normalized observation fails to capture dynamic structures such as business cycles.
- Macroeconomic states: GAN without macroeconomic variables has an out-of-sample Sharpe ratio around 10% lower than GAN with hidden macroeconomic states.The comparison supports using extracted states rather than the full macroeconomic panel as predictors.
- Macroeconomic states: The GAN model with hidden states outperforms the unconditional model by about 20% in Sharpe ratio.The unconditional model includes LSTM states in SDF weights but uses a constant conditioning function.
D. Predictive Performance
The model’s risk loadings predict future stock returns and generate an almost linear security market line, although benchmark factor models fail to explain the resulting return differences.
- Higher β portfolios have higher subsequent returns, with the highest and lowest deciles clearly separating.This indicates that the model’s risk loadings predict future stock returns.
- β-sorted portfolio returns are almost perfectly linear in average β, with R2 values of 0.98, 0.97, and 0.95 across three sorting resolutions.The reported values correspond to quintiles, deciles, and 20 quantiles, respectively.
- The fitted security market line has an intercept slightly below zero, indicating a very good but not perfect fit.
- Market and Fama-French factors do not explain the systematic return differences across β-sorted portfolios.
- The GRS test rejects correct pricing by the market and Fama-French factor models for the β-sorted portfolios.
E. Pricing of Characteristic Sorted Portfolios
The model achieves strong pricing performance across characteristic-sorted portfolios, with its advantage concentrated in momentum- and reversal-related portfolios and nonlinear interactions.
- GAN substantially better captures variation and mean returns for short-term reversal and momentum decile portfolios than EN and FFN.
- GAN’s advantage is driven by extreme portfolios, especially the tenth short-term-reversal decile and first momentum decile.
- Book-to-market and size portfolios are easy to price, with all models achieving time-series R2 above 70% and cross-sectional R2 close to 1.
- GAN captures nonlinear interactions that linear models miss, explaining its stronger performance on double-sorted reversal and momentum portfolios.
- The findings generalize to other decile-sorted portfolios and to value-weighted portfolios.
F. Variable Importance
Variable-importance analysis identifies trading-friction and past-return characteristics as central inputs, while learned macroeconomic states track business-cycle conditions.
- All three models select trading frictions and past returns as the most relevant characteristic categories.
- For GAN, the most important characteristics are short-term reversal, standard unexplained volume, and momentum.
- GAN’s first 20 variables represent all six major anomaly categories, including value, intangibles, investment, and profitability.
- The two macroeconomic variables that stand out are the median bid-ask spread and the federal funds rate.They are interpreted as capturing overall economic activity and market volatility.
- The four hidden macroeconomic states exhibit cyclical behavior and peak, particularly for the third and fourth states, during recessions.
- The states do not coincide at all times, indicating that they capture different macroeconomic information.
G. SDF Structure
Individual characteristics often affect the stochastic discount factor approximately linearly, but GAN’s nonlinear interactions across characteristics explain its stronger performance.
- Individual characteristics have an almost linear effect on the pricing kernel and risk loadings in one-dimensional relationships.
- GAN’s stronger performance is explained by nonlinear interaction effects that capture dependencies between multiple characteristics.
- GAN exhibits nonlinearities around the median that correspond to decile portfolios where it outperforms FFN and EN.
- GAN allows low and high characteristic quantiles to have different linear slopes, unlike a purely additive parallel-shift structure.
- GAN reveals more complex interaction patterns, including distinct exposure to value for small versus large stocks.
- The size and book-to-market characteristics have a highly nonlinear joint effect on the GAN pricing kernel.
H. Robustness Results
The GAN model remains robust across stock-size samples, tuning parameters, estimation windows, and limits-to-arbitrage controls. Its performance and economic structure remain strong for large-cap stocks and across alternative specifications.
- Large-capitalization stocks: GAN’s explained variation is two to three times higher than for the linear or deep learning prediction models.The cross-sectional R2 gap is also substantially wider on larger stocks than in the full sample.
- Large-capitalization stocks: GAN identifies systematic structure in large-cap stocks, whereas FFN and linear models mainly fit small stocks.When estimated and evaluated on stocks above the 0.001% market-capitalization threshold, GAN performance is essentially identical, suggesting the same SDF structure is recovered.
- Tuning parameters: Alternative GAN fits have essentially identical asset-pricing performance and SDF correlations above 80%.The alternative models differ in network depth and the number of instruments constructing test assets, yet the authors report replicable pricing performance and economic models.
- Estimation window: A rolling-window GAN SDF has 70% correlation with the benchmark and similar general patterns, with no major improvements from time-varying re-estimation.The authors conclude that results are robust to the estimation time window.
- Limits to arbitrage: Removing trading-friction and past-return characteristics leaves pricing performance unchanged and produces a 78% correlation with the benchmark SDF.This robustness test addresses whether GAN is targeting illiquid-stock anomalies rather than economic risk.
I. Machine Learning Investment
The paper argues that machine-learning signal extraction and portfolio construction should be integrated through the stochastic discount factor. The resulting GAN portfolio retains high risk-adjusted returns while allowing explicit trade-offs with trading frictions.
- Portfolio performance: The GAN factor has comparable or lower turnover than the other SDF portfolios, suggesting similar exposure to transaction costs.The SDF is presented as a tradeable portfolio with an attractive risk-return trade-off.
- Portfolio performance: GAN’s normalized cumulative return exceeds those of the other models while avoiding fluctuations and large losses.The paper reports that its Sharpe ratio is by far the highest and its drawdown and maximum loss are comparable to the other models.
- Trading frictions: Trading-friction screening reveals a clear trade-off between reduced frictions and achievable Sharpe ratios.Stocks below specified market-capitalization, bid-ask-spread, or turnover quantiles are removed by setting their SDF weights to zero.
- Trading frictions: 1.73, 2.07, and 1.87 annual Sharpe ratios remain after removing 40% of the smallest stocks, highest bid-ask spreads, and least-active stocks, respectively.These are lower bounds because the GAN was not re-estimated after removing the stocks.
- Machine-learning portfolio design: The SDF is the conditionally mean-variance efficient portfolio, motivating joint signal extraction and portfolio design.The approach targets signals most relevant for overall portfolio design rather than predicting returns in a separate first step.
J. SDF of Multi-Factor Models
The paper evaluates how GAN-based stochastic discount factors complement multi-factor models by combining factor structure with flexible, conditional pricing weights. IPCA combined with GAN performs strongly out-of-sample across Sharpe ratios, explained variation, and cross-sectional pricing performance.
- Comparison with IPCA: IPCA factors use more information than the paper’s models, which may contribute to their high Sharpe ratios.The comparison therefore reflects different information sets.
- One-factor reductions: The IPCA one-factor model can achieve a negative cross-sectional R2, while Sharpe-ratio maximization reduces explained variation to around 1%.Optimizing one objective does not preserve performance on other asset-pricing dimensions.
- Nonlinear loadings: Flexible nonlinear loadings explain more variation and produce smaller pricing errors than linear loadings.This comparison concerns alternative SDF specifications within the IPCA framework.
- GAN and IPCA results: IPCA combined with GAN is the best-performing model across the reported dimensions, with high Sharpe ratios, near-benchmark XS-R2, and among the highest explained variation.Its Sharpe ratios exceed the benchmark GAN for K ≥7, while only the fully flexible benchmark GAN has better explained variation.
- Interpretation: The GAN framework is complementary to multi-factor models because it uses factor-imposed structure together with additional information in the factors.The paper presents this complementarity as applying the GAN framework to structured multi-factor representations.
- Model design: The model uses no-arbitrage constraints, flexible conditioning, and time-varying macroeconomic information to estimate stock-return pricing relationships.Its outputs include risk measures β and SDF weights ω as functions of characteristics and portfolio weights.
- Macroeconomic information: Macroeconomic data can be uninformative for asset pricing when only their latest increments are used as inputs.The paper emphasizes accounting for the time dimension of financial data.
Appendix C. Implementation
The implementation combines adaptive optimization, dropout regularization, hyperparameter selection, and a purpose-built adversarial network to estimate conditional SDFs. The procedure solves successive optimization problems to construct instruments and price instrumented returns, while simulations assess the roles of no-arbitrage, flexibility, and macroeconomic dynamics.
- Adam adapts the learning rate, helping optimization escape saddle points and converge faster.
- Dropout regularizes training by temporarily removing network units, while retaining all units during out-of-sample testing.
- The procedure searches 384 hyperparameter combinations, selects four validation winners, re-estimates each with nine initializations, and retains the best ensemble.
- The conditional network nonlinearly transforms firm characteristics and hidden macroeconomic states before combining them linearly to generate moments.
- The adversarial estimation sequentially prices 10,000 stocks, generates an 8-dimensional instrument vector, and prices 80,000 instrumented stock returns.
- Simulations indicate that no-arbitrage, flexible interactions, and LSTM-based macroeconomic dynamics are needed, whereas forecasting and simple linear formulations cannot achieve these goals.
1. Two characteristics: The loadings are the multiplicative interaction of two characteristics
The simulations examine nonlinear characteristic interactions and time-varying macroeconomic states. GAN outperforms forecasting and linear alternatives because it captures interaction structure, hidden business-cycle dynamics, and the associated conditional SDF behavior.
- Two characteristics: The two-characteristic setup produces nonlinear interaction effects and low signal-to-noise ratios for many assets.
- Two characteristics: The GAN model outperforms forecasting and linear models across all reported evaluation categories in the two-characteristic simulation.
- Two characteristics: The GAN reaches the population SDF’s Sharpe ratio, while the forecasting approach attains a high Sharpe ratio but fails to explain systematic variation.
- Two characteristics: GAN assigns positive weights to high/high and low/low characteristic combinations and negative weights to high/low and low/high combinations.
- Macroeconomic state variable: First differences appear stationary but lose business-cycle information; the GAN’s LSTM instead extracts the hidden macroeconomic state process.
- Macroeconomic state variable: The GAN strongly outperforms forecasting and linear models in the macroeconomic-state simulation, where the linear model uses the wrong SDF-weight sign roughly half the time.
- Simulation conclusions: The simulations show that recent macroeconomic observations alone rule out general dynamics, while no-arbitrage helps address low signal-to-noise ratios.
- Adversarial approach: The adversarial approach is designed to improve robustness to misspecification, address weak pricing factors, identify SDF parameters, and allow time-varying moments.
Appendix E. Asset Pricing Results for Sorted Portfolios
The appendix reports out-of-sample evaluations across characteristic-sorted portfolios, alternative GAN specifications, risk measures, turnover, variable importance, and rolling-window estimates. It documents the portfolio designs and model configurations used for these comparisons.
- Sorted portfolios: The appendix evaluates explained variation and pricing errors for portfolios sorted by short-term reversal, momentum, size, and book-to-market ratio.
- Variable importance: Variable-importance analyses rank 178 macroeconomic variables and 46 firm-specific characteristics using normalized average absolute gradients on test data.
- Factor comparisons: The appendix reports GAN SDF correlations and time-series regressions against the five Fama-French factors, including monthly SDF pricing errors.
- Portfolio risks: Risk comparisons report Sharpe ratios, maximum one-month losses, and maximum drawdowns, while turnover is reported separately for long and short positions.
- Model specifications: The benchmark GAN uses two 64-node SDF layers, four hidden macroeconomic states, eight instruments, and 32 hidden macroeconomic states in the adversarial network.
- Model specifications: Alternative GAN estimations include architectures with 4 layers and 32 instruments, 2 layers and 8 instruments, or 3–4 layers and 16 instruments.
- Rolling-window estimates: Rolling-window variable importance and SDF weights are averaged across estimates re-estimated on 240-month windows.