Noisy Financial Machine Learning
Professor Federico Bandi
James Carey Endowed Professor
Johns Hopkins University, Carey Business School
Even in a noisy environment – such as the cross-sectional prediction of individual stock returns – the traditional benchmark of no predictability in the machine learning literature appears to be excessively conservative should the objective be operational cross-sectional predictions/rankings. Outperforming the no predictability benchmark may, in fact, just amount to some degree of market timing. Consistent with this observation, we document that the panel historical mean constitutes a more revealing benchmark, one which permits clearer separation between market-timing (i.e., level) effects and genuine cross-sectional predictions/rankings. Not only is the panel historical mean economically justifiable (as the mean of the equally-weighted market portfolio), it is also encompassed by typical machine learning models as the case in which the characteristics play no role. We formalize its predictive role in a finite-sample theory for Ridge under small signal-to-noise ratio. When tested against the panel historical mean, only rich machine learning models yield meaningful predictions. They, however, do so for smaller, value, less liquid stocks.














