Why your quant model knew everything and still lost money


In 1847, Ignaz Semmelweis showed that mortality rates in Vienna General Hospital’s maternity ward dropped 90% when doctors washed their hands with chlorinated lime solution between the morgue and the delivery room. The correlation was unambiguous. His colleagues rejected the finding. The reason they rejected it is important. He could show what happened, but he could not explain why. Without a causal mechanism — germ theory would not come for another decade — the correlation looked like an artefact. How could clean hands have anything to do with childbirth mortality? Semmelweis died in an asylum in 1865, discredited. Pasteur proved he was right fifteen years later.

Modern quantitative finance has the opposite problem. It generates correlations at industrial scale. Every factor, every signal, every cross-asset relationship is measured to six decimal places. What it systematically fails to do is ask Semmelweis’s question: why does this relationship exist, and will it exist tomorrow?

The correlation trap

In August 2007, quant funds collectively lost between $25 billion and $30 billion in a period now referred to as the quant quake. The factors that had generated alpha for a decade unwound simultaneously at a speed that correlation matrices had no mechanism to predict. The funds had not made modelling errors. Their correlations were accurate. What they had modelled was not the causal structure of markets — it was the historical co-movement of markets. These are different things, and they diverge precisely at the moments when you most need them to agree.

This is the correlation trap. Train a model on historical data; it learns to predict in the historical distribution. Introduce a regime shift — a liquidity crisis, a geopolitical shock, a central bank pivot — and the distribution changes. The causal structure of the market does not change: prices are still driven by supply, demand, information asymmetry, and risk appetite. But the historical correlations that proxied for that structure no longer hold. The model knows everything. It just doesn’t understand any of it.

What causal AI actually means

Causal machine learning is the application of Judea Pearl’s causal hierarchy to the problem of financial signal generation. Pearl, who won the Turing Award in 2018, established three levels of reasoning about data:

  • Association — what correlates with what? (standard ML)
  • Intervention — what would happen if I changed X? (causal inference)
  • Counterfactual — why did this happen, and what would have happened differently?

Standard machine learning operates entirely at Level 1. It can tell you that momentum in EUR/USD over the past 20 days correlates with the next 5-day return. It cannot tell you whether that relationship is causal, coincidental, or the result of a confounding variable that may not persist. Marcos López de Prado, whose work on financial machine learning has become standard reading at systematic funds, put it directly in Causal Factor Investing (Cambridge University Press, 2023): “Most factors in the academic literature are spurious. They are statistical artefacts that survived multiple testing in a particular historical period. Causal factors — those driven by structural economic mechanisms — are rare and far more valuable.”

The most accessible causal inference method for practitioners is Double Machine Learning

The practical application: Double Machine Learning

The most accessible causal inference method for practitioners is Double Machine Learning (DML), introduced by Chernozhukov et al. (2018) at MIT. The idea is to partial out the confounding effect of control variables using ML, then estimate the causal effect of the variable of interest on the residuals.

In a trading context: you want to know whether a specific microstructure signal — say, order flow imbalance — causally predicts short-term returns, or whether its apparent predictive power is an artefact of correlated macro variables (VIX level, time of day, news flow). DML partials out those confounders using two separate ML models, then estimates the causal relationship in the residuals. The resulting estimate is far more likely to generalise out-of-sample because it strips away the spurious historical correlations. Consider a concrete example. An analyst observes that EUR/USD order flow imbalance over 30-minute windows has predicted 15-minute forward returns with an information coefficient of 0.08 over a two-year sample. Promising. But that sample included the 2022 EUR/USD trending regime — elevated directional flow, persistent VIX spikes, concentrated macroeconomic data release clustering — all of which correlate with both order flow and short-term returns. DML fits one model predicting order flow imbalance from those confounders, and a second model predicting returns from the same confounders. The residual-on-residual regression strips away the macro environment and isolates what the order flow signal actually explains independently. The IC is smaller. But it is structural — and it holds out of sample, through regime changes, because the causal mechanism is invariant.

Bryan Kelly at Yale’s School of Management, whose work on ML for asset pricing has been widely cited, has argued: “The next frontier in empirical asset pricing is not building better prediction machines, but understanding the economic mechanisms that generate predictable variation in returns.” Causal methods are that frontier.

Which FX signals are genuinely causal?

The causal filter divides FX signals into two categories: those driven by structural economic mechanisms that persist across regimes, and those that are statistical artefacts of a particular historical period. Three signals illustrate the distinction clearly. 

Order flow imbalance. When informed traders know something the market does not, they trade directionally before the information is publicly revealed. The resulting order flow imbalance does not merely correlate with price movement — it causes it, because informed demand forces market-makers to widen spreads and adjust quotes to compensate for adverse selection risk. VPIN (Volume-Synchronized Probability of Informed Trading), developed by Easley, López de Prado, and O’Hara at Cornell, quantifies this structural mechanism directly. The causal pathway is invariant across regimes: informed flow ® adverse selection ® quote adjustment ® price impact. It does not depend on the prevailing macro environment. It depends on the existence of informed trading, which is a permanent feature of financial markets.

For FX derivatives structuring desks, this mechanism raises a harder, largely untested question. Illiquid option prints are sparse and disjointed — sometimes weeks or months apart for a given tenor and structure — which is exactly why the standard microstructure toolkit has never been applied to them. But if informed spot flow genuinely drives adverse selection, does that toxicity transmit forward into the market impact of an illiquid derivatives execution — or is post-trade spread widening on those contracts driven by something else entirely: dealer inventory, vol surface stress, structural hedging flow? That is not a rhetorical question. It is precisely the kind of hypothesis DML exists to test — one DeepAlgo Sovereign is now putting to that same scrutiny, rather than assuming the answer either way.

Interest rate differentials. Carry — buying high-yield currencies, selling low-yield — reflects a structural constraint: uncovered interest parity. Rational agents should equalise risk-adjusted returns across currencies after adjusting for exchange rate expectations. That UIP fails systematically in practice is not a statistical artefact; it reflects a structural feature of risk appetite, funding constraints, and investor heterogeneity. The mechanism is invariant; only its magnitude changes across economic cycles. A causal model of carry does not rely on the EUR/USD-specific correlation history of 2015-2023. It relies on the structural economics of interest rate differentials, which are regime-independent.

Price momentum. Here, causal scrutiny becomes uncomfortable. Momentum has generated positive returns across four decades of backtests, yet its mechanism remains genuinely contested — underreaction, risk compensation, or crowding from trend-following capital. A signal whose mechanism is disputed is precisely the one that collapses during regime shifts, because there is no principled way to know whether it has disappeared or merely changed form.

The practical implication is a portfolio construction principle: signals with identified structural mechanisms warrant higher base conviction and more stable position sizing. Signals that are historically correlated but causally unvalidated should carry wider confidence bands and be revalidated more frequently as market regimes evolve.

FX markets are uniquely challenging for standard ML because the regime distribution is highly non-stationary

Why this matters specifically for FX

FX markets are uniquely challenging for standard ML because the regime distribution is highly non-stationary. A model trained on EUR/USD dynamics during 2021-2023 — elevated volatility, strong directional trends, aggressive central bank divergence — has learnt correlations from a specific macro environment. As that environment normalised through 2024-2025, many of those correlations weakened or reversed.

A causal model trained to identify structural mechanisms degrades more slowly precisely because the mechanisms themselves do not change with the regime. The relationship between order flow imbalance and adverse selection is as operative in a low-volatility consolidating market as it is in a trending crisis. The carry mechanism functions regardless of whether the cycle is in a risk-on or risk-off phase. Only the parameters — the magnitude of the effect — change.

That durability is the commercial argument for causal approaches in FX. Not that causal models are more accurate in-sample — they often are not. But that they fail more slowly and more predictably out-of-sample, which is the environment where the money is actually made and lost.

The honest caveat

Causal inference in finance is hard. A significant challenge is that financial markets do not permit controlled experiments — the fundamental tool of causal identification. Pearl himself has noted that causal inference from observational data requires “causal assumptions that cannot be verified from the data alone.”

For financial practitioners, this means causal methods should be used to validate hypotheses that have economic intuition — not as a mechanical replacement for standard ML. The goal is not to replace your factor model. It is to stress-test it: does this signal survive causal scrutiny, or does its predictive power collapse when you partial out the confounders?

Conclusion

Semmelweis had the right answer. He just could not explain why. For two decades, financial ML has had the right data. What it has lacked is the right question. The question is not: what predicts returns in my training set? The question is: what causes return variation in a way that will persist when the market environment changes?

That distinction is the difference between a model that performs in backtest and a model that performs in production. Causal AI is how you close the gap. And in FX markets — where regime shifts arrive without warning and historical correlations can reverse overnight — closing that gap is not an academic exercise. It is a survival requirement.

Charles Glah, ASIP is Founder and Chief Product Officer of DeepAlgo Sovereign and Bletchley Intelligence Ltd. He specialises in applied machine learning for systematic trading, regime detection, FX derivatives structuring, and AI governance.

We will be happy to hear your thoughts

Leave a reply

Som2ny Network
Logo
Compare items
  • Total (0)
Compare
0
Shopping cart