Most machine learning alpha in systematic trading doesn't decay because the signal was arbitraged away—it decays because the model was never capturing alpha in the first place, just overfitting to transient liquidity microstructure artifacts that practitioners quietly mistake for edge because their backtests can't distinguish the two.
Most ML alpha doesn't die because someone found your signal and traded it away. It dies because it was never real alpha to begin with. What practitioners mourn as "signal decay" is usually the quiet expiration of overfit artifacts masquerading as edge, and our collective inability to distinguish the two is the most expensive blind spot in quantitative finance.
The numbers are damning. The median live Sharpe ratio of ML driven systematic funds falls 40 to 70 percent below backtest Sharpe within the first 12 to 18 months of deployment. Marcos López de Prado's Probability of Backtest Overfitting framework demonstrates the mechanism clearly: given enough trials, nearly all backtested strategies appear profitable. This is not a theoretical curiosity. It is the default outcome of standard research workflows. Yet most practitioners who experience this degradation reach for a comforting explanation rather than the correct one.
The dominant narrative goes like this: alpha decays because markets are adaptive. You discover a profitable signal, deploy capital against it, competitors notice, and competing capital arbitrages the edge away until it vanishes. This is the "alpha lifecycle" framework popularized by AQR and taught in virtually every quant program. It dominates conference panels at QuantCon and Battle of the Quants. JPMorgan's QDS group publishes factor crowding metrics that implicitly reinforce this worldview. The narrative is intellectually comfortable because it preserves the practitioner's belief that the model was correct. It simply got crowded out. Your alpha was real. You were just too slow, or too small, or too late.
This framing, while valid for a narrow set of well known factors, is dangerously incomplete. It ignores a far more prevalent failure mode that nobody wants to talk about honestly.
Here is what actually happens in the majority of cases. Models trained on historical data latch onto transient liquidity microstructure artifacts. Fleeting bid/ask bounce patterns. Temporary market maker inventory effects. Time of day volume clustering anomalies. Exchange specific queue priority behaviors. These artifacts appear statistically robust in sample. They produce impressive Sharpe ratios, clean equity curves, and all the visual signatures of genuine edge. But they have no persistent economic mechanism behind them. They are phantoms of a particular microstructure regime, not expressions of a durable informational advantage.
Standard backtesting infrastructure cannot distinguish these artifacts from genuine alpha because both produce identical statistical signatures in finite samples. A naive ML model trained on tick or minute bar data will happily learn spread capture patterns that are entirely unrealizable net of transaction costs and latency, yet show spectacular backtest performance. When practitioners then tune hyperparameters across hundreds of model specifications using overlapping data windows, they are compounding the problem at industrial scale. Harvey and Liu's 2019 Journal of Financial Economics paper on the factor zoo demonstrated that most "discovered" factors fail proper multiple testing adjustments. ML hyperparameter search is the same statistical mirage, just automated and accelerated. The model isn't discovering edge. It is memorizing the noise fingerprint of a specific microstructure environment that will not persist.
When you forensically decompose the return streams of ML strategies that experienced sharp post deployment degradation, the culprit is overwhelmingly microstructure contamination rather than competitive arbitrage. The tell is in the decay signature itself. True crowding driven alpha decay is gradual. It correlates with observable increases in competitor AUM and rising strategy correlation across the industry. Artifact decay looks completely different. It is sudden, regime independent, and uncorrelated with any external crowding metric.
I have seen this pattern repeatedly. A mid frequency equities strategy built on an ensemble gradient boosted model shows a 3.2 backtest Sharpe. It collapses to 0.4 live. Forensic decomposition reveals that over 60 percent of the backtest PnL originated from sub second price patterns tied to exchange specific queue priority artifacts on a single venue. That venue then updated its matching engine logic. The "alpha" vanished overnight. Contrast this with genuine crowding decay as documented in AQR's paper "How Can a Strategy Still Be Valuable If Everyone Knows About It?" where value and momentum factors decayed slowly over decades, proportional to institutional adoption. These are categorically different phenomena, yet the industry labels both "alpha decay."
Sophisticated practitioners defend against this failure mode not through better models but through more adversarial validation infrastructure. The discipline is architectural, not algorithmic. It lives in the pipeline, not in model selection. This means transaction cost models calibrated to live fill data with conservative slippage assumptions. It means purged and embargoed walk forward cross validation as López de Prado prescribes to prevent train test leakage. It means a mandatory narrative for every signal: who is the counterparty and why are they systematically wrong? It means synthetic data perturbation tests that deliberately alter microstructure properties to confirm the signal survives the change. There is a reason firms like Two Sigma and Citadel Securities are known to invest more engineering resources in execution simulation than in alpha research itself. They learned this lesson with real money.
If the quant industry honestly audited how much of its historical "alpha decay" was actually artifact expiration, the implications would reshape fund due diligence, allocator evaluation frameworks, and the entire economics of systematic strategy development. Much of the capital allocated to systematic strategies was never buying real edge. It was renting statistical noise with a convincing backtest attached.
López de Prado estimates over 90 percent of backtested strategies are false discoveries. Apply the decay signature diagnostic to your own book. If your worst drawdowns are uncorrelated with measurable increases in strategy crowding or competitor deployment, the parsimonious explanation is not that someone stole your alpha. It is that the alpha was a ghost. So the uncomfortable question is not whether your alpha decayed. It is whether you can prove it was ever there.