Recent interesting research from Cakici and Zaremba, highlights an often-overlooked aspect of machine learning for equity return prediction: the choice of prediction target. Rather than focusing on increasingly sophisticated model architectures or feature engineering, the authors show that how returns are represented during training has a much larger impact on predictive performance. In particular, models trained to predict stock ranks instead of raw return levels generate substantially stronger portfolio performance—roughly doubling both returns and Sharpe ratios in large-cap universes.
The evidence comes from a comprehensive study of more than 80,000 stocks across 35 global equity markets spanning 1994–2024. The authors systematically compare different transformations of both input characteristics and target returns, finding that preprocessing the target variable is the single most important design choice for improving cross-sectional return forecasts. While the benefits of rank-based targets are pronounced in large-cap stocks, the results vary across firm sizes and market environments, suggesting that no single transformation is universally optimal.
The core mechanism reflects a trade-off between robustness and information preservation. Rank-based targets (percentile mappings, rank-to-[−1,1], Gaussianized ranks) filter noise and align naturally with portfolio sorting, boosting average monthly alphas from ~1.0% (raw targets) to ~1.9%. However, by discarding magnitude information, these transformations underperform when return distributions exhibit high dispersion or skewness—environments where extreme outcomes carry economically meaningful signals. In such settings, magnitude-preserving transformations (demeaning, standardization) allow models to map extreme signals to extreme outcomes more effectively, particularly among small- and micro-cap stocks where tail behavior dominates cross-sectional variation.
This heterogeneity extends across borders: no single transformation dominates globally. Rank-based targets excel in stable, developed markets with tighter return distributions, whereas standardized targets outperform in emerging markets with higher volatility and asymmetry. For practitioners, this implies that the choice of transformation should be adaptive—conditioned on the market regime, firm-size exposure, and distributional properties—rather than fixed. The study also reveals that while dynamic selection based on past performance yields modest gains (~30–50 bps monthly), the instability of relative performance over time limits the efficacy of simple adaptive rules.
Figure 1 illustrates that target transformations drive alpha improvements far exceeding those from feature preprocessing; Figure 2 confirms consistent global coverage across the 35-market sample; Table 3 shows rank-based targets delivering ~1.6–1.9% monthly returns versus ~0.9–1.1% for raw targets; and Table 5 reveals that magnitude-preserving transformations dominate in micro-cap segments where rank-based methods lose their edge.
Authors: Nusret Cakici and Adam Zaremba
Title: Getting the Target Right in Return Prediction
Link: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6615698
Abstract:
We show that a largely overlooked design choice-how returns are defined as prediction targets-drives machine learning performance in stock returns. Using a large international panel of equities from 35 markets over 1994 to 2024, we compare transformations applied to stock characteristics and target returns. Transforming the target from raw to standardized or rank-based returns nearly triples predictive accuracy and doubles portfolio returns. Feature transformations play a secondary role. Rank-based targets perform best on average but discard information about return magnitudes, leading to underperformance when dispersion or skewness is high, particularly among micro-cap stocks. The optimal transformation varies across markets and over time.
As continually, we present several interesting figures and tables:




Notable quotations from the academic research paper:
“[…] transformations of stock characteristics and target returns affect machine learning performance in asset pricing. Using a large international sample spanning 35 markets from 1994 to 2024, we systematically vary how both characteristics and returns are constructed. We apply a unified set of transformations—standardization, winsorization, nonlinear scaling, and rank mappings—to both sides of the prediction problem. We estimate expected returns with an ensemble of machine learning models and evaluate the implied portfolio strategies. This design isolates how data representation affects predictive accuracy and economic outcomes.
The data reveals a systematic pattern behind this trade-off. The relative performance of transformations varies with the shape of the return distribution. In segments characterized by high dispersion or pronounced skewness, magnitude-preserving transformations dominate. These environments feature large cross-sectional differences and asymmetric tail behavior, where extreme outcomes carry economically meaningful information. Rank-based methods compress this variation and mask differences across the tails, reducing predictive performance. This effect appears consistently across settings: among small and micro-cap stocks, in equal-weighted portfolios that tilt toward such firms, and in markets with volatile, dispersed, and skewed returns.
Cross-country evidence substantiates this mechanism. No single transformation dominates globally, and heterogeneity is markedly larger for targets than for features. Rank-based targets perform best in more stable, developed markets, where return distributions are tighter and less skewed. In contrast, standardized and other magnitude- preserving transformations perform better in markets with higher dispersion and more pronounced asymmetries, especially emerging markets. Country-level tests confirm that dispersion and skewness are key determinants of relative transformation performance, linking the cross-country evidence directly to the underlying shape of the return distribution.
[The] paper shows that how returns are defined as prediction targets has large effects on machine learning performance in asset pricing. While the literature has focused on model choice and feature engineering, our evidence indicates that how the target is represented is also of key importance. Transforming returns—through demeaning, standardization, or rank mapping—substantially improves both predictive accuracy and portfolio performance. Rank-based targets perform best on average, reflecting their robustness to outliers and their alignment with portfolio construction. However, this advantage is not universal.
[…] findings imply that data representation, especially the definition of the prediction target, is a central modeling decision rather than a preprocessing detail. Focusing on models while holding the target fixed risks optimizing the wrong margin. Effective return prediction requires adapting the representation of returns to the economic environment rather than relying on a single specification.”
Are you looking for more strategies to read about? Sign up for our newsletter or visit our Blog or Screener.
Do you want to learn more about Quantpedia Premium service? Check how Quantpedia works, our mission and Premium pricing offer.
Do you want to learn more about Quantpedia Pro service? Check its description, watch videos, review reporting capabilities and visit our pricing offer.
Are you looking for historical data or backtesting platforms? Check our list of Algo Trading Discounts.
Or follow us on:
Facebook Group, Facebook Page, Twitter, Linkedin, Medium or Youtube
Share onLinkedInTwitterFacebookRefer to a friend