explorational · a divergent move

Many
models

opensframe bymultiplying frompsychology & decision research withalone riskduplicates in disguise

Any single model of a messy question is wrong somewhere. What do you gain by running several different models at once and treating their disagreement as the finding rather than a nuisance?

Scott Page calls it many-model thinking: instead of betting a decision on one clever lens, you point several different lenses at the same question and read them together. Each lens is wrong in its own way — a straight-line trend can only ever draw a straight line; a mean-reverting model always expects a return to the average; a random walk refuses to predict anything but yesterday. Run them side by side and their forecasts fan out into a bracket, and the width of that bracket — the place where models disagree — is not noise to be tidied away. It is the single best picture you have of your own uncertainty. Page's diversity-prediction theorem makes the promise exact: the ensemble's error is always the average model's error minus the diversity among the models. Disagreement is what buys the accuracy.

There is a catch, and it is this move's named risk. The theorem rewards diversity, not count. Bolt together four models that secretly share the same assumption and they will agree loudly with one another — a narrow, confident bracket — while being wrong together, because the diversity term has gone to zero. This is duplicates in disguise: three near-identical linear fits dressed up as an ensemble, offering the warm feeling of consensus and none of its substance. The instrument below lets you build a real ensemble and watch it beat its own best member out of sample — and then lets you stack near-copies and watch a comfortable, meaningless agreement form in front of you.

the signature instrument

One quantity, several honest forecasts.

A hidden process generates the history you can see (dots) and a future you cannot (the dashed truth is revealed for scoring). Each model fits only the visible history and forecasts forward; the thick teal line is their average. The displayed values — per-model error, ensemble error, the diversity decomposition, the bracket width — are computed live from the model on held-out points.

The bracket

4 diverse models · diminishing-returns process · horizon 12

Models in the ensemble

The hidden true process

Observation noise · σ

2.5

Forecast horizon · held-out points

12

Duplication · how alike the near-copies are

0.00

Held-out error · root-mean-square, lower is better

The diversity-prediction decomposition · mean squared error on held-out points

Ensemble error
=
Avg individual error
Diversity

the bracket is the finding

Disagreement, read as information.

The instinct with several forecasts is to pick the best one, or to split the difference and move on. Many-model thinking asks you to do something stranger: to look first at where the models part company. A single model hands you a point and a false serenity — one number, no visible doubt. Four different models hand you a fan, and the fan is honest. When the linear trend, the mean-reversion, the exponential, and the random walk all cluster tight, the future they describe is genuinely constrained and you may act with some confidence. When they blow apart, that width is telling you the truth that any single model was hiding: this question is underdetermined, and the data you have cannot yet close it.

The move therefore opens the frame by multiplying. Each model is a full, formal lens — a complete little theory of how the quantity behaves — and no lens is neutral. Hold several at once so that the question is bracketed rather than answered prematurely. In the instrument, drag the horizon out and watch the bracket widen: honest uncertainty grows with distance, exactly as it should, and only an ensemble shows you that growth. The single confident line does not get less certain as it marches into the future. It just gets more wrong, silently.

what to try

Three experiments on the instrument above.

01

Start on the bracket. Read the error bars: the ensemble's root-mean-square error sits below every single model, including the best of them. Average four flawed forecasts and the average is better than the best one you could have picked — because their errors point in different directions and partly cancel. Reseed a few times; the ensemble keeps winning.

02

Now hit false agreement. The ensemble is three near-identical linear fits. The bracket collapses to a sliver, the picture looks reassuringly tight — and the diversity term falls almost to zero, so the ensemble barely improves on its members. A narrow bracket among similar models is not confidence. It is the same model, counted three times.

03

Switch the process to regime shift: the world trends up, then turns down past the divider. Watch the linear and exponential models sail off in the wrong direction while the ensemble bends and degrades gracefully. No single lens survives the turn — but the crowd of them loses far less than the worst of them.

diversity, not count, is the gain

Where the accuracy actually comes from.

The decomposition under the chart is not a metaphor; it is an identity that holds exactly, on every click. For squared error, the ensemble's error equals the average model's error minus the diversity of the models:

ensemble error  =  average individual error  −  diversity

The average individual error is how good your models are on their own. The diversity is how much they disagree with each other — the mean squared distance from each model's forecast to the ensemble's. Since diversity can only ever be positive, the ensemble is never worse than the average member, and it is better by exactly the amount they diverge. You do not get accuracy by adding more models; you get it by adding different ones. Two brilliant forecasters who always agree are worth barely more than one; two mediocre ones who err in opposite directions can beat them both. When the instrument's diversity term is large, the ensemble pulls far ahead. When you stack duplicates and drive diversity to zero, the ensemble sinks back to the level of any one of them — the arithmetic refuses to reward the redundancy.

The same identity imposes an honest limit: the theorem guarantees the ensemble beats the average member, not necessarily the best member you could have chosen with hindsight. In the default it does beat the best — but flip to a sharp regime shift and one lucky model can occasionally edge it out. The ensemble is the robust bet across regimes, not a law that it wins every single hand.

model

How the measures work.

A hidden process generates a signal — a trend, a diminishing-returns curve, runaway exponential growth, a mean-reverting wave, or a regime that turns — and the visible history is that signal plus Gaussian observation noise, drawn from a seeded generator so every run is reproducible. Four structurally different models each fit only the 40 visible points: the linear model by least squares; the mean-reversion by decaying from the last value back toward the historical average; the exponential by a least-squares fit in log-space; the random walk by simply holding the last value flat. Each then forecasts the held-out horizon.

The errors are scored against the underlying signal at the held-out points — the structure, not the fresh noise — so the comparison isolates what each lens gets structurally right. Per-model error is root-mean-square; the decomposition is reported in mean-squared units, where the identity ensemble error = average error − diversity holds to machine precision. The near-duplicate models are the linear fit with its slope nudged by a small amount that the duplication dial shrinks toward zero; at full duplication they become exact copies and diversity vanishes. Treat the instrument as a disciplined ensemble you can break on purpose — not an oracle. The models here are deliberately simple; the lesson about diversity is what transfers, not the particular four lenses.

the move ↔ the instrument

What each part stands for.

The moveThe instrument
a modelone formal lens on the question — a complete little theory, wrong in its own particular way.
the forecast bracketthe span of honest uncertainty: where genuinely different models refuse to agree.
the ensemble averagethe combined estimate — the single number that survives all the lenses at once.
the diversity termhow much disagreement buys accuracy; the exact amount the ensemble beats its average member by.
near-duplicate modelsfalse agreement — a narrow, comforting bracket that is really one model under several names.
a regime shiftthe world no single lens survives; the test of whether the crowd degrades gracefully.

how this opening fails

The failure modes to hold in view.

risk · duplicates in disguise

Several models that share one assumption.

The catalogue's named risk for this move. An ensemble looks rigorous, but rigour is an illusion if the members are one idea in different fonts — three linear fits, or four models that all assume the trend continues, will produce a narrow bracket and agree themselves into confident error. The diversity term is the tell: when it collapses toward zero, the agreement is structural, not evidential. Test each pair by asking whether any future could make them disagree. If nothing could, they are the same model, and the ensemble is counting one witness as four.

risk · the average can smear a sharp signal

Averaging is a default, not a law.

When one model is right and the rest are noise, the ensemble drags the good forecast toward the bad ones — averaging away a correct sharp signal into a blurred compromise. The diversity-prediction theorem guarantees you beat the average member, never that you beat the best member, and under a clean regime a single well-matched model can win outright. So the ensemble is the right bet under genuine uncertainty about which lens fits — precisely because you rarely know in advance which model is the right one — but it is a disciplined default, not a substitute for knowing. When you truly do know, use the model you trust; when you do not, let the crowd of them bracket your ignorance.

Run several different models; the width where they disagree is the truth about your uncertainty.

can you use it?

Three questions before you go.

RECOGNITION — Which of these is the move? A: fitting a linear trend, a mean-reverter, and a random walk to the same history and reading their spread. B: running the linear fit in three different software packages and citing the consensus. C: asking three colleagues to guess the number and averaging.

Answer

A. Structurally different lenses, disagreement read together. B is the near-miss: duplicates in disguise — one model in three fonts, diversity zero, agreement structural rather than evidential.

THE NEAREST NEIGHBOR — Splitting the difference between forecasts also yields one combined number. What does many-model thinking read that the compromiser throws away?

Answer

The width. Where the models part company is the finding — tight means the future is constrained, wide means the data cannot yet close the question. Averaging first discards your best picture of the uncertainty.

PRODUCTION — Budget season: you must forecast next year's support-ticket volume, and two dueling spreadsheets have deadlocked the room. Run the move before looking.

One version + the check

Fit three genuinely different lenses — straight trend, diminishing returns, mean reversion — forecast with each, and present the bracket, budgeting against its width. Yours works if some possible future would separate every pair of your models, and the width itself made it into the recommendation.