Research · Essay
When Agents Disagree
We do not build one model that knows. We build several that argue, and we listen to the argument.
By NDI Research ✦ ✦3 minutes’ reading
In Brief
- The system is composed of specialised agents with different data, different methods and, by design, different incentives.
- Averaging independent forecasts reduces error only if their errors are independent. Most ensembles quietly fail this test.
- Disagreement between agents is not noise to be removed. It is a direct measurement of how much the problem is actually understood.
The most natural way to build an intelligent system is to make it a single, very capable model, give it all available information and ask it what will happen. This produces fluent answers. It also produces a system with a single point of view, a single set of blind spots and no internal mechanism for noticing either.
We build differently. The system is a set of agents, each with a narrow responsibility and a deliberately limited view of the problem. Some observe. Some forecast. One exists only to argue against whatever the others conclude. Intelligence, to the extent the system has any, emerges from the friction between them.
§ IRoles
| Agent | Responsibility | Deliberately does not see |
|---|---|---|
| Observer | Collects and structures market and external information; maintains the evidence base | Other agents' forecasts |
| Forecaster | Produces probabilities from structural and flow-based models (several independent instances) | Market-implied probabilities |
| Historian | Finds comparable historical regimes and reports base rates | Current narratives and commentary |
| Contrarian | Constructs the strongest case against the emerging consensus of the other agents | Its own previous positions |
| Risk | Attempts to invalidate any proposed decision before it is made | Expected return |
| Allocator | Converts aggregated probabilities into exposure under constraints | Individual agent identities |
| Executor | Handles programmable execution and reports realised costs | The forecast rationale |
| Auditor | Scores past decisions, measures calibration, adjusts weights | Nothing. The Auditor sees everything, after the fact |
The third column matters as much as the second. Independence between forecasts is not a property you can add after the fact; it has to be engineered by controlling what each agent is allowed to know. The Forecaster instances never see market-implied probabilities, because a model that sees the answer will drift towards it. The Historian never sees current commentary, because commentary is where narratives live.
§ IIThe independence problem
The case for ensembles rests on a simple statistical fact: the average of several forecasts has lower expected error than the typical individual forecast, provided their errors are not perfectly correlated. The benefit grows as correlation falls. With five forecasters whose errors are correlated at 0.9, the ensemble is barely better than any one of them. At 0.3, it is substantially better.1
Most ensembles in practice are built from models trained on the same data, tuned against the same benchmarks and shaped by the same assumptions. Their errors are highly correlated, and their agreement is mistaken for confidence. We monitor pairwise error correlation between agents continuously. When it rises above a threshold, the Auditor treats the affected agents as a single forecaster for weighting purposes until diversity is restored.
§ IIIAggregation
Agent forecasts are combined in log-odds space rather than averaged as probabilities. A weighted mean of log-odds respects the fact that moving from 90% to 99% is a much larger claim than moving from 50% to 59%. Weights are set by each agent's recent calibration on comparable questions.
| Agent | Weight | Forecast |
|---|---|---|
| Forecaster · structural | 0.30 | 78% |
| Forecaster · flows | 0.25 | 66% |
| Historian | 0.25 | 61% |
| Contrarian | 0.20 | 38% |
| Pooled (α = 1.25) | 67% | |
| External reference | 44% |
§ IVDisagreement as a measurement
In the example above, the agents span forty points. That spread is not discarded once the pooled probability is computed. It is recorded alongside it as a dispersion measure — the weighted standard deviation of agent log-odds, 0.61 in this case — and it flows directly into confidence and sizing.
Two questions can carry the same pooled probability of 67% and deserve very different treatment. If every agent independently arrives near 67%, the problem is well understood, and the remaining uncertainty is genuine randomness. If the agents range from 38% to 78%, the problem is poorly understood, and the 67% is an average of incompatible models of the world. The first is a forecast. The second is a disagreement with a number attached.
“When our agents agree, we ask whether they are independent. When they disagree, we ask which of them understands something the others do not.”
§ VThe Contrarian is not an opposite
The Contrarian agent is often misunderstood, including, early on, by us. It does not simply invert the consensus of the other agents. Its task is to construct the most plausible coherent world in which the emerging view is wrong, and to estimate how likely that world is. When it cannot construct one, its forecast converges with the others, and that convergence is itself informative.
In practice, the Contrarian is the agent most often wrong and the agent whose removal most damages ensemble calibration. It loses arguments frequently. It prevents the others from becoming certain.