explorational · a divergent move

Degrees of
belief

openscommitment bysuspending frompsychology & decision research withalone riskflattening

A yes-or-no verdict deletes its own remainder — the thirty percent you were wrong about vanishes the instant you round to "true." What do you gain by holding beliefs as movable probabilities instead?

Philip Tetlock spent two decades scoring the predictions of people paid to be certain, and the finding that made superforecasting a word was almost embarrassingly plain: the best forecasters do not know more than the confident experts they beat. They think differently about what they know. They report granular, oddly specific credences — sixty-eight percent, not "likely" — and they revise them in small, frequent steps as evidence trickles in. A superforecaster who says seventy keeps the thirty alive as a working scenario, a version of the world still on the table, still funded with a little attention. The expert who rounds to "it will happen" has thrown that thirty away, and with it any way of noticing, later, that it was the part that mattered.

Binary thinking feels decisive, and decisiveness is flattering; but a verdict is a point where a belief is a region, and the rounding that turns one into the other is not free. It discards information, and — more — it punishes calibration, because a scoring rule that grades your zeros and ones has no way to reward the person who honestly said "seventy" and was right seven times in ten. The instrument below runs three forecasters over the very same stream of resolvable events and lets you watch, in live numbers, what the rounding costs. Then it hands you the slider and scores you.

the signature instrument

A calibration & Brier-score trainer.

Three forecasters call the same stream of yes/no events. The granular one reports a credence near each event's hidden true probability; the binary one rounds every credence to 0 or 100; the fence-sitter says 0.5 to everything. Every number — each Brier score, each calibration point — is computed from the resolved outcomes, not asserted. Drag the controls, reseed, or flip on play mode and forecast the stream yourself.

The same events, three ways of believing

300 rounds · resolved live · Brier = mean (p − outcome)²

Rounds resolved300
how many resolvable events the forecasters call
Forecaster noise · skill0.10
how far the granular credence strays from true p · lower = sharper skill
Event difficulty0.30
higher = events closer to a coin flip · true p pulled toward 0.5
Play it yourself 
your own Brier and calibration build alongside the three
granular · near true p binary · rounds to 0 / 100 fence-sitter · 0.5 always

Brier score — lower is better

a perfect forecast scores 0 · a coin-flip guess on balanced events scores 0.25

Calibration — do your 70%s happen 70% of the time?

the deleted remainder

What a binary verdict throws away.

Round the granular forecaster's seventy up to a hundred and you have, in effect, promised the event will happen. When it does, you look magnificent — you were "right," fully. When it doesn't, the scoring rule collects: a confident 1.0 against a NO outcome is the maximum possible error, four times worse than an honest coin-flip. The binary forecaster is not making a different bet than the granular one; it is making the same bet with the hedge stripped out. It keeps all the upside of being right and volunteers for all the downside of being wrong, and over a long enough stream that trade is a loser. In the instrument it loses every time: across a hundred random seeds at the default settings, the granular credence has a strictly lower Brier than the binary rounder in a hundred of them.

The thirty percent deleted by the binary verdict justified keeping a fallback warm, phrasing the email cautiously, and not betting the rent. Degrees of belief are just the refusal to delete it — to carry the belief as a number that still remembers how wrong it might be. That is the divergent move here: it opens commitment by suspending the reflex that collapses a probability into a verdict before the world has actually resolved. You do not decline to act; you act while keeping the remainder legible.

what to try

Three moves on the trainer above.

01

Flip on play it yourself and forecast a dozen events on the slider. Watch your calibration curve bend off the diagonal exactly where you were overconfident — where your stated 80%s came true only 60% of the time, the point sags below the line.

02

While playing, read the tally line: it shows your Brier and what it would have been had you rounded every forecast to 0 or 100. Watch the rounded number jump upward. That gap is the price of the deleted remainder, charged to you personally.

03

Load the fence-sitter preset. Its Brier settles near 0.25 and its calibration collapses to a single dot — perfectly hedged, and perfectly useless. Then raise difficulty toward a coin flip and watch even the reckless binary pundit's Brier climb past it.

calibration is a skill you can see

Reading the curve and the score together.

A single forecast is unfalsifiable — if you say seventy and it happens, were you right? You cannot know from one event. Calibration is what a stream reveals. Gather every time a forecaster said "about seventy," look at how often those events actually came true, and plot the pair. Do it for every band of confidence and you get the calibration curve. A forecaster whose seventies happen seventy percent of the time, whose nineties happen ninety percent, lands on the diagonal — their stated confidence means exactly what it says. The granular forecaster in the instrument hugs that line, wobbling only from the sampling noise of finitely many events. The binary forecaster cannot: having only ever said 0 or 100, it can plot only two points, marooned at the corners, and its "hundreds" come true nowhere near a hundred percent of the time.

The Brier score is the same virtue compressed to one number: the mean squared distance between what you said and what happened. It rewards two things at once. One is calibration — meaning what you say. The other is sharpness: how far you dare to stray from a limp 0.5. The fence-sitter is beautifully calibrated in a trivial sense — on balanced events its lone 0.5 point sits right on the diagonal — but it has zero sharpness, so it never helps anyone decide anything. Confidence is only worth claiming if it is earned: extreme credences pay when they are calibrated and cost when they are not. The granular forecaster wins not by being bold and not by being cautious, but by being honest about how much it actually knows.

the move ↔ the trainer

What each part of the instrument stands for.

The moveThe trainer
a credencea belief held as a probability — a number on the slider, not a verdict, still remembering its own remainder.
rounding to 0 / 1the binary verdict that deletes the remainder: all the upside of being right, all the downside of being wrong.
the Brier scorethe cost of being confidently wrong, tallied over a whole stream — mean squared error between credence and outcome.
the calibration curvewhether your seventies actually happen seventy percent of the time; distance from the diagonal is self-deception, measured.
sharpnessearned confidence — how far you stray from 0.5, which only pays if the calibration is there to back it.
0.5 for everythingflattening: hedging as a refusal to think, calibrated in the trivial sense and useless in every other.

model

How the measures work.

No faked data. Each round draws a hidden true probability p from a distribution centred on 0.5, whose spread the difficulty dial controls — high difficulty squeezes every p toward a coin flip. The event then resolves YES with probability exactly p and NO otherwise, using a seeded generator so a given seed replays the same stream. The granular forecaster reports p plus Gaussian noise (the noise dial is that spread; at zero it reports the truth and is perfectly calibrated by construction). The binary forecaster rounds that credence to 1 if it exceeds 0.5, else 0. The fence-sitter reports 0.5 flat.

Scoring is disclosed and plain. The Brier score is the mean over resolved rounds of (credence − outcome)², where outcome is 1 for YES and 0 for NO; it runs from 0 (perfect) to 1 (perfectly, confidently wrong). A forecaster who says 0.5 to everything scores exactly 0.25 on every round because (0.5 − 0)² and (0.5 − 1)² are both 0.25. A forecaster reporting the true p cannot beat the irreducible average of p(1 − p), the world's own coin-flip variance. Hard streams therefore give everyone a higher Brier floor. The calibration curve bins forecasts into ten intervals of stated probability, and for each populated bin plots the mean stated probability against the observed YES frequency; with enough events and a calibrated forecaster it recovers the diagonal to within sampling error. Treat the whole thing as a practice range, not a crystal ball: the events are simulated so their truth is knowable and the scoring exact.

how this opening fails

The failure modes to hold in view.

risk · flattening

Degrees of belief degenerate into a permanent 0.5.

The catalogue's named risk for this move. Held wrongly, "everything is a probability" becomes an excuse never to commit to a number that costs anything — every question answered with a noncommittal fifty-fifty, every view held "equally," none engaged. The fence-sitter in the instrument is exactly this failure, and its low-ish Brier is a trap: it is calibrated in the trivial sense while carrying zero sharpness, which is to say it never helps anyone decide. Degrees of belief are the opposite discipline — the willingness to say the number, seventy and not fifty, and to be scored on it. A credence you would never stake anything on is not a belief; it is a way of not having one.

risk · needs a stream

Brier and calibration need many resolvable events.

Both the score and the curve are properties of a stream, not of a single call. They are cheap and sharp for weather, sports, and repeated operational forecasts — and nearly silent for the calls we most want help with: the rare, unrepeatable, once-in-a-career decision that resolves exactly once, if ever. You can hold a credence about whether a war starts or a marriage lasts, but you cannot calibrate it, because there is no bin of similar events to check it against. The method is weakest precisely where certainty is most seductive. Hold the number honestly there too — just without the false comfort that the arithmetic has vouched for it.

Say the number; the remainder you keep alive is the part that turns out to matter.

can you use it?

Three questions before you go.

RECOGNITION — Which is holding degrees of belief? A: 'I'm not sure, let's wait.' B: 'I'm at seventy percent this ships on time — the thirty is the vendor, so I'm watching them.' C: 'It'll definitely ship.'

Answer

B. A movable percentage keeps its remainder alive and pointed at what would move it. A suspends entirely; C collapses to certainty and deletes the remainder.

THE NEAREST NEIGHBOR — Suspension of judgment also resists premature closure. What separates degrees of belief from it?

Answer

Commitment. Suspension declines to hold a view at all; degrees of belief commits partially and acts on it — the seventy is enough to plan around while the thirty stays funded and watched.

PRODUCTION — Take a claim you'd normally state flatly ('the launch will succeed'). Put a number on it, then name what specific observation would move you five points either way.

One version + the check

A version: '65%. A strong first-week retention number moves me to 75; churn above 8% drops me to 55.' Yours works if the number has a named trigger for updating, not a static hedge.