explorational · a divergent move

Competing
hypotheses

openshypotheses bymultiplying fromstrategy & institutions within a group riskduplicates in disguise

How do you keep evidence — not attachment — from choosing your explanation? The moment you adopt one story, it becomes an intellectual child you defend. Analysis of Competing Hypotheses (Heuer, CIA 1999) forces every explanation onto the table at once and scores the evidence against all of them, so the winner is the story the evidence fails to kill.

Thomas Chamberlin saw the trap in 1890. The instant the mind adopts a single explanation, he wrote, "affection for his intellectual child springs into existence, and as the explanation grows into a definite theory his parental affections cluster about his offspring." From then on the work inverts: you stop asking whether the theory is true and start gathering reasons it might be. Every ambiguous fact is read in its favour. Contradictions are filed as anomalies to be explained later. The single hypothesis does not lose arguments — it recruits them.

Richard Heuer's answer, built for intelligence analysts after a run of confident, wrong assessments, is deliberately counterintuitive: you make progress not by confirming your favourite but by disconfirming the rest. Put every rival explanation on the table at once. Run each piece of evidence across all of them and ask a colder question than "does this fit my theory?": "which theories does this rule out?" The move exposes a quiet swindle in ordinary reasoning: the most seductive evidence is often the most worthless, because a fact consistent with every story tells you nothing about which story is true. The instrument below lets you watch that happen — and refuse to be moved by it.

the signature instrument

An honest competing-hypotheses matrix.

Click any cell to cycle its rating — Consistent → Neutral → Inconsistent. Drag a row's reliability to weaken shaky evidence; toggle a row off to remove it. The displayed values — posteriors, diagnosticity, inconsistency — are computed live from a real naive-Bayes model over your ratings.

Chest pain in the clinic

6 evidence rows · 4 hypotheses · likelihoods 0.9 / 0.5 / 0.1

Consistent · P(E|H)=0.9 Neutral · 0.5 Inconsistent · 0.1 ◆ non-diagnostic row = fits every story

Posterior belief — where the evidence lands

Heuer's ranking — fewest inconsistencies

    the intellectual child

    Why multiplying is the move.

    A single hypothesis is not a neutral starting point; it is a gravitational one. Once it exists, it bends the search for evidence toward itself, because the cheapest cognitive operation in the world is finding another reason to believe what you already believe. You can spend a week confirming a wrong theory and feel, the entire time, that you are being rigorous — because you are gathering data, running tests, reading closely. The rigour is real. It is just pointed the wrong way.

    The competing-hypotheses move breaks the gravity by refusing to let any one explanation stand alone. It opens the hypotheses dimension by multiplying: instead of "is it a heart attack?" — a yes/no that already smuggles in the answer you fear — you write down heart attack, reflux, anxiety, and strain as equals, and only then bring in the evidence. The matrix is a discipline for holding several intellectual children at once so that you cannot love any of them too early. Loyalty shifts from a favoured answer to the inquiry itself.

    what to try

    Three experiments on the matrix above.

    01

    Find "Recent stressful period." The model flags it non-diagnostic — in the preset it's Consistent with every hypothesis, so its diagnosticity is zero. Toggle it off and on. The posterior bars do not move. A fact that feels like it matters, that fits your story perfectly, changes nothing — because it fits every rival just as well.

    02

    Take "ECG shows ST changes" — currently Consistent with heart attack, Inconsistent with the rest — and drag its reliability down toward zero. Watch the heart-attack bar deflate. One "damning" piece of evidence loses its entire grip the moment you stop trusting it. Certainty and content both affect the result.

    03

    Flip a single cell from Consistent to Inconsistent under the leading hypothesis. One disconfirmation drops it hard — harder than several confirmations lifted it. That asymmetry is Heuer's whole thesis: evidence against is worth more than evidence for.

    disconfirmation does the work

    Why a fact consistent with everything proves nothing.

    Heuer's ranking rule is not "the hypothesis with the most check-marks wins." It is the opposite: rank by the fewest inconsistencies. His reasoning is that consistency is cheap — most hypotheses are consistent with most evidence, which is exactly why analysts drown in confirmations and still get it wrong. Inconsistency is expensive and therefore informative: a single genuine contradiction can eliminate a hypothesis that a dozen agreeable facts had propped up. So the analyst's attention belongs on the rows that discriminate — the ones where the ratings differ sharply across hypotheses.

    That is what the diagnosticity column measures: the spread of a row's likelihoods across the hypotheses. A row that reads Consistent for one story and Inconsistent for another has high diagnosticity: it discriminates. A row that reads the same for all of them has diagnosticity near zero, and the instrument flags it: consistent with every story, so it moves nothing. The evidence that feels most damning is frequently the evidence that discriminates least, precisely because a vivid, salient fact tends to fit whatever narrative you bring to it.

    The two rankings in the instrument can disagree. On the shipped preset, the naive-Bayes posterior favours heart attack, while Heuer's fewest-inconsistencies rule favours anxiety — because anxiety, though weakly supported, is contradicted by only one row. The tool surfaces both on purpose. A method that always agreed with itself would be hiding a choice from you; this one shows you that "most probable given a model" and "hardest to rule out" are different questions, and lets you decide which you trust.

    model

    How the measures work.

    Each rating becomes a likelihood P(evidence | hypothesis): Consistent = 0.9, Neutral = 0.5, Inconsistent = 0.1. Each row carries a reliability weight w that shrinks its likelihood toward the uninformative middle: P = 0.5 + w·(base − 0.5). At w = 1 the rating is taken at face value; at w = 0 the row is inert, worth exactly 0.5 for every hypothesis — evidence you don't trust cannot vote. Priors default to uniform (you can edit each hypothesis's prior in its column header).

    The posterior for each hypothesis is its prior times the product of its likelihoods over the rows that are switched on, normalised so the column sums to 1: P(H) ∝ prior · Πⱼ P(Eⱼ | H). That "product of likelihoods" step is the naive-Bayes independence assumption — it treats every piece of evidence as an independent witness. It is a deliberate simplification, and a consequential one: if two rows are really saying the same thing, the model counts them twice. Diagnosticity is the plain spread of a row's likelihoods, max minus min, across the hypotheses. All of it recomputes on every click. Treat the output as a disciplined scratchpad that makes your own ratings argue with each other — not an oracle that knows the answer.

    the move ↔ the matrix

    What each part of the instrument stands for.

    The moveThe matrix
    a hypothesis columna rival explanation forced onto the table as an equal, before any evidence is weighed.
    a Consistent / Inconsistent ratingevidence for or against the world under that particular story — a likelihood, not a verdict.
    the posteriorwhere belief actually lands once every switched-on row has been scored against every hypothesis.
    diagnosticityhow sharply a piece of evidence discriminates — how differently the rival stories predict it.
    a non-diagnostic rowthe seductive fact that fits every story equally, and therefore moves the posterior not at all.
    the inconsistency scoreHeuer's disconfirmation ranking: the hypothesis hardest to rule out, not the one most agreed with.

    how this opening fails

    The failure modes to hold in view.

    risk · duplicates in disguise

    Four hypotheses that are really one.

    The catalogue's named risk for this move. A full matrix looks rigorous, but rigour is an illusion if the columns are one idea under four names — "sabotage," "insider threat," "deliberate act," and "not an accident" are not four hypotheses, they are one. Padding the table with near-duplicates gives every real rival less room and flatters whatever theme they share. The test is whether a single piece of evidence could be Consistent with one column and Inconsistent with its neighbour. If no evidence can ever separate two columns, they are the same hypothesis, and the matrix is rigged.

    risk · the model is a simplification

    Naive-Bayes double-counts, and the likelihoods are fixed.

    The independence assumption treats every row as a separate witness, but real evidence clusters: "ECG shows ST changes" and "radiates to left arm" may both be downstream of the same cardiac event, and the model happily multiplies them as if they were unrelated, over-driving the posterior. The 0.9 / 0.5 / 0.1 mapping is a convenient fiction too — few facts are truly nine-to-one. So the exact percentages are softer than they look. What the instrument gets right is structural, not numerical: which evidence discriminates, which is inert, and how much a disconfirmation costs. Read it as a scratchpad that disciplines your judgement, never as a verdict that replaces it.

    Put every story on the table, and keep the one the evidence cannot kill.

    can you use it?

    Three questions before you go.

    RECOGNITION — Which is competing-hypotheses analysis? A: brainstorming every explanation for the outage. B: listing explanations, then scoring each piece of evidence against every one and keeping the evidence that discriminates. C: picking the likeliest cause and gathering support for it.

    Answer

    B. The method is the matrix: evidence scored against all hypotheses, so you can see which evidence separates them. A only generates; C confirms a single story.

    THE NEAREST NEIGHBOR — Multiple working hypotheses also holds several explanations at once. What does competing-hypotheses analysis add?

    Answer

    Diagnosticity. Multiple working hypotheses keeps the family alive; competing-hypotheses analysis scores the evidence and surfaces which items are consistent with everything — and therefore prove nothing — moving attention to what actually discriminates.

    PRODUCTION — Sales dropped last quarter. List three hypotheses, then name one piece of evidence that is consistent with all three (so proves nothing) and one that would separate them.

    One version + the check

    A version: hypotheses — pricing, a competitor, seasonality. 'Traffic was down' fits all three. 'Down only in regions the competitor entered' discriminates. Yours works if you found at least one non-diagnostic item and one that splits the field.