Factor InvestingPortfolio Theory

Machine Learning Lifts Stock Portfolio Sharpe Ratio to 2.4

Summary by Robert Gorak · Published September 4, 2026 · Last reviewed September 4, 2026

Bryan T. Kelly and Dacheng Xiu·2023·Foundations and Trends in Finance

Financial machine learning applies penalized regression, decision trees, and neural networks to return prediction, factor pricing, and portfolio choice. Kelly and Xiu (2023) survey this literature in Financial Machine Learning, highlighting Gu, Kelly, and Xiu (2020)'s model comparison. Tested on the CRSP-Compustat panel over the 1987-2016 test period, monthly R² rises from -3.46% for OLS to 0.40% for a neural network. Kelly, Pruitt, and Su (2020) separately price systematic risk in 11,452 US stocks using almost 95% fewer parameters than standard factor models.

What the Study Found

Gu, Kelly, and Xiu (2020) find out-of-sample monthly R² for predicting stock returns of -3.46% for OLS versus 0.40% for a neural network (NN3). Long-short decile portfolios sorted on neural network forecasts earn annualized equal-weight Sharpe ratios of 2.1 (NN1), 2.4 (NN3), and 2.2 (NN5). A three-signal linear benchmark earns only a 0.8 Sharpe ratio over the same period. Kelly, Pruitt, and Su (2020) price systematic risk in stock returns using almost 95% fewer parameters than observable factor models. Simon, Weibels, and Zimmermann (2022) find that adding a neural network to a portfolio rule lifts its Sharpe ratio from 1.8 to 2.5.

"A primary goal of our work is to help readers recognize machine learning as an indispensable tool for developing our understanding of financial markets phenomena."

Kelly and Xiu (2023), Financial Machine Learning.

Methodology

Kelly and Xiu (2023) synthesize the machine learning literature on return prediction, factor models, discount factor estimation, and portfolio choice. Its central benchmark, from Gu, Kelly, and Xiu (2020), evaluates thirteen models on the CRSP-Compustat panel of US stock-month returns. Models train on 1957-1974 data, validate on 1975-1986 data, and are tested out-of-sample over 1987-2016. The survey's own theoretical contribution uses ridge regression and random matrix theory rather than a new empirical sample.

Key Statistics

Metric Finding Context
Out-of-sample monthly R² (OLS) -3.46% Full US stock panel, test period 1987-2016 (Gu, Kelly & Xiu, 2020)
Out-of-sample monthly R² (NN3, neural network) 0.40% Full US stock panel, test period 1987-2016 (Gu, Kelly & Xiu, 2020)
Long-short decile Sharpe ratio (NN3) 2.4 Equal-weighted, annualized (Gu, Kelly & Xiu, 2020)
IPCA parameter reduction vs. observable factor models ~95% fewer parameters 11,452 stocks, 37 instruments, 599 months (Kelly, Pruitt & Su, 2020)
Maximum Sharpe Ratio Regression (MSRR) 1 = β'Ft + ut Recasts mean-variance portfolio choice as OLS regression on characteristic-managed factors
Elastic net penalty RSS + λ(1-ρ)Σ|βj| + ½λρΣβj² Penalized loss underlying regularized linear return-prediction models

OLS vs Neural Network Return Predictions

Measure OLS Neural Network (NN3)
Out-of-sample monthly R² (all stocks) -3.46% 0.40%
Out-of-sample monthly R² (top 1,000 stocks) -11.28% 0.70%
Out-of-sample monthly R² (bottom 1,000 stocks) -1.30% 0.45%

Why This Matters

Parsimonious models are not automatically safer than heavily parameterized ones once regularization is applied correctly. Quantitative asset managers can allocate research budgets toward nonlinear methods and larger feature sets instead of hand-picked linear factors. Researchers designing new prediction models should evaluate out-of-sample Sharpe ratios directly, rather than treating a small R² as proof of little economic value. A principled account of why complexity helps also gives portfolio managers a basis for choosing model size beyond intuition about overfitting.

Frequently Asked Questions

Financial machine learning uses penalized regression, decision trees, and neural networks that lift out-of-sample stock-return R² to 0.40%, versus -3.46% for ordinary least squares. Kelly and Xiu (2023) survey these methods across return prediction, factor pricing, discount factor estimation, and portfolio construction.

Neural networks reach an out-of-sample monthly R² of 0.40% versus -3.46% for OLS, across the full US stock panel over 1987-2016. Gu, Kelly, and Xiu (2020) document this margin, with the network's R² rising to 0.70% among the top 1,000 stocks by market value.

Long-short decile portfolios sorted on neural network return forecasts earn annualized equal-weight Sharpe ratios of 2.1 to 2.4, versus 0.8 for a three-signal linear benchmark. Adding a neural network to a portfolio rule raised one strategy's Sharpe ratio from 1.8 to 2.5 (Simon, Weibels & Zimmermann, 2022).

Kelly, Pruitt, and Su (2020) use almost 95% fewer parameters than observable factor models, across a panel of 11,452 US stocks. Their sample spans 37 instruments over 599 months. Their Instrumented Principal Components Analysis matches those models' descriptive fit, including the Fama-French five-factor model.

Reference

Bryan T. Kelly and Dacheng Xiu (2023). Financial Machine Learning. Foundations and Trends in Finance.

Read the full paper

Cite this summary

Gorak, R. (2026). Machine Learning Lifts Stock Portfolio Sharpe Ratio to 2.4. Tradicted. https://www.tradicted.com/research/kelly-machine-2023/