Research · Essay
Models Fail Before Narratives Do
On regime change, model decay and the long half-life of a good story.
By NDI Research ✦ ✦4 minutes’ reading
In Brief
- Regime changes are usually visible first in the behaviour of models — rising error, falling agreement — and only much later in the stories people tell about markets.
- A model that has stopped working does not announce it. It continues to produce confident numbers. Decay has to be detected from the outside.
- We monitor three early signals of decay. In July, one of them fired eighteen days before our outlook failed. We did not act on it.
Narratives are durable. A story about why markets behave the way they do — liquidity is abundant, policy is supportive, volatility is structurally suppressed — can survive months of contradicting evidence, because it is maintained by people who have built positions, careers and frameworks around it. Models have no such loyalty. When the structure they were fitted to changes, they degrade, often quickly, and the degradation shows up in their errors long before anyone changes the story.
This asymmetry is an opportunity and a hazard. It is an opportunity because model behaviour can be an early indicator of regime change. It is a hazard because a decaying model does not know it is decaying. It produces the same confident probabilities, in the same format, until it is shown to be wrong.
§ IThree signals
We monitor three properties of the ensemble continuously. None is a forecast; each is a measurement of whether the forecasting machinery is still connected to the world.
- Rolling calibration error. The distance between stated probabilities and observed frequencies over a trailing window, including interim questions that resolve weekly. A rise indicates that the models' sense of likelihood has drifted.
- Residual structure. In a working model, errors should be unpredictable. When errors begin to correlate with each other or with an observable variable, the model is systematically missing something.
- Agreement collapse. Independent agents that have historically agreed begin to diverge. A sharp fall in model agreement often means the agents are responding to a change that some of their methods capture and others do not.
- All agents
- Liquidity-sensitive agents
§ IIJuly, again
The July failure is described in Calibration Before Conviction. What that note does not say is that the agreement signal fired on 13 July. Agreement among liquidity-sensitive agents fell from 79% at the end of June to 51% in the second week of July, crossing the decay threshold. Under the rules then in place, a decay flag lowered confidence on new questions but did not alter forecasts already issued for the month.
In retrospect, the flag was the system telling us, in its own terms, that the liquidity regime it had been fitted to was ending. The narrative — liquidity is improving — remained intact in external commentary for most of August. Our models had abandoned it within two weeks.
Since August, a decay flag on a model family triggers an interim revision of every open question that depends on it, published on the system page. We expect this rule to generate false alarms. We prefer false alarms to a repeat of July.
§ IIIDecay is normal
It is tempting to treat model failure as an exception — a bug to be fixed, after which the model will work again. We have come to think of it as the default condition. Every model is fitted to a particular structure of the world, and every structure eventually changes. The relevant question is never whether a model will fail, but how quickly its failure will be detected and how much damage it will do in the meantime.
| Model family | Fitted regime | Median useful life | Current status |
|---|---|---|---|
| Funding and liquidity | Post-2023 balance-sheet runoff | 14 months | Rebuilt, Aug 2026 |
| Systematic positioning | Vol-targeting mechanics | > 36 months | Active |
| Energy supply | Spare-capacity regime | 9 months | Down-weighted |
| Policy path | Data-dependent easing | 11 months | Active |
The table carries an implication we find uncomfortable. The models that last longest are mechanical: they describe the arithmetic of how certain participants must behave given their rules, not the beliefs of anyone. The models that decay fastest are the ones that encode a view about the world — precisely the ones that generate the most interesting disagreements with consensus.
§ IVHolding views loosely
The practical consequence is a preference for detecting failure over avoiding it. We do not believe we can build models that do not decay. We believe we can build a system that notices decay faster than the stories around it change, and that withdraws confidence before it withdraws capital. That, more than any single forecast, is what we mean by intelligence that does not default.