Judgement, Calibration and Tradecraft

Foresight lives or dies on the quality of the judgement behind it. This module examines why unaided expert judgement about the future is systematically unreliable, and what disciplines exist to correct for it. It moves from the psychology of prediction through structured analytic techniques, the US Intelligence Community's analytic standards, reference-class forecasting, and the accountability that comes from scoring forecasts after the fact. Throughout it draws on the tradecraft canon (Richards Heuer, ICD 203, Kahneman and Tversky, Tetlock, Flyvbjerg) and on how DSI Advisory applies that canon to the security industry's own alarms, most directly in its Alarm Test and its Predictions Scorecard.

  • analytic-tradecraft
  • competing-hypotheses
  • calibration
  • icd-203
  • reference-class-forecasting
  • structured-analytic-techniques
13 min · Core

The Psychology of Prediction

Before any method can help, an analyst has to accept an uncomfortable finding: unaided expert judgement about the future is unreliable, and confidence is a poor guide to accuracy. This lesson sets out the biases that corrupt prediction, why expertise does not immunise against them, and why the rest of the module treats tradecraft as a corrective rather than a decoration.

~4 min

By the end you can

  • Name the principal cognitive biases that distort judgement about the future.
  • Explain why domain expertise does not remove those biases and can deepen them.
  • Describe how group dynamics amplify individual error.
  • Justify why structured method is needed, not optional, for serious foresight.

Confidence is not accuracy

The most durable finding in the study of judgement is that people, including experts, are far more confident about the future than their record justifies. Daniel Kahneman and Amos Tversky spent decades showing that the mind does not weigh evidence the way a careful accountant would. It reaches for whatever answer comes to hand quickly and then dresses that answer in the language of reason. For an analyst whose product is a forecast, this is not an academic curiosity. It means the felt sense of certainty that accompanies a judgement carries almost no information about whether the judgement is right.

The biases that corrupt a forecast

Several distortions recur. Anchoring pulls an estimate toward whatever number was mentioned first, even an irrelevant one. Availability makes vivid, recent or heavily reported events feel more probable than they are, which is why the last spectacular breach shapes a threat assessment more than the quiet, common failure. Hindsight bias convinces us after the fact that we knew all along, which quietly destroys the feedback an analyst needs to improve. Confirmation bias is the most corrosive of all: once a favoured explanation forms, the mind hunts for evidence that supports it and discounts evidence that does not. Heuer, writing for the intelligence community, called this the central problem of analysis, because the analyst who loves a hypothesis will find reasons to keep it long after the facts have turned.

Why expertise does not save you

It is tempting to assume that experience inoculates the specialist. It does not. Expertise sharpens pattern recognition, which is precisely what makes confirmation bias more dangerous: the expert sees the familiar pattern faster and commits to it sooner. Philip Tetlock's long study of political forecasters found that the most famous and most confident experts were often the least accurate, because a strong prior theory made them explain away every disconfirming signal. Knowing more can mean being wrong with greater authority.

When the group makes it worse

Analysis rarely happens alone, and the group can amplify rather than cancel individual error. Groupthink, Irving Janis's term, describes how a cohesive team suppresses dissent to preserve harmony, converging on a comfortable consensus and mistaking agreement for validation. A room that never surfaces a serious counter-argument has not reached certainty; it has reached silence. The intelligence failures reviewed after major surprises repeatedly show this pattern: not an absence of information, but an absence of challenge.

Why method, not willpower

The natural response is to resolve to try harder, to be more objective. This does not work, because the biases operate beneath awareness and resist introspection. You cannot see your own blind spot by staring harder. What corrects the failure is external structure: procedures that force the consideration of alternatives, that make reasoning visible, and that hold a forecast to account after the fact. That is the argument for everything that follows. Tradecraft is not decoration on top of good judgement; it is the scaffolding that makes good judgement possible where instinct alone fails.

Unaided expert judgement fails because these biases resist introspection, not effort.
Unaided expert judgement fails because these biases resist introspection, not effort.

Check your understanding

Answer each from memory. Your results are saved in this browser and count toward your readiness — sign in (account panel above) to keep them across devices.

  1. Why is an expert's felt confidence a poor guide to whether a forecast is correct?

  2. Which of these best describes why domain expertise does not remove confirmation bias?

  3. What does groupthink, in Irving Janis's sense, do to a team's analysis?

14 min · Core

Analysis of Competing Hypotheses

If confirmation bias is the central failure, Analysis of Competing Hypotheses is the central correction. Richards Heuer's technique inverts the natural habit: instead of building a case for a favoured answer, the analyst lays out every plausible hypothesis and tries to disprove each one. This lesson explains the method, its logic of falsification, and how DSI's Alarm Test turns the same discipline back on the security industry's own claims.

~4 min

By the end you can

  • Describe the steps and logic of Analysis of Competing Hypotheses.
  • Explain why disproving hypotheses is stronger than confirming a favourite.
  • Relate ACH to Popperian falsification and structured analytic techniques.
  • Explain how DSI's Alarm Test applies ACH to the security industry's own alarms.

Inverting the natural habit

Left to itself, the mind picks a likely answer and then gathers support for it. Richards Heuer, in Psychology of Intelligence Analysis, argued that this is exactly backwards, and designed a procedure to reverse it. Analysis of Competing HypothesesRichards Heuer's structured technique that lists all plausible hypotheses and assesses evidence against them simultaneously, favouring the hypothesis with the fewest inconsistencies rather than the most support., or ACH, begins by listing every reasonable explanation for a situation, including the ones the analyst finds unattractive. Only then is evidence brought in, and it is assessed one item at a time against all the hypotheses at once, in a matrix. The question for each piece of evidence is not whether it fits the preferred story but which hypotheses it is inconsistent with.

Disproof beats confirmation

The power of the method comes from that shift. Evidence consistent with a hypothesis is weak, because it is usually consistent with several hypotheses at the same time; the fact that a country is buying weapons is consistent with both aggressive and defensive intent. What discriminates is diagnostic evidence, evidence that is inconsistent with some hypotheses. So ACH scores hypotheses by their inconsistencies and asks which explanation has the fewest pieces of evidence weighing against it, rather than which has the most in favour. You do not select the winner; you eliminate the losers. This is why an analyst can hold a well-supported hypothesis and still be asked, correctly, what would have to be true for it to be wrong.

The Popperian root

This is Karl Popper's philosophy of science made operational. Popper held that a claim earns credibility not by accumulating confirmations, which are easy to find, but by surviving serious attempts to falsify it. A hypothesis that cannot in principle be disproved is not strong; it is empty. ACH is one of a family of structured analytic techniques, alongside devil's advocacy, key-assumptions checks and red-teaming, that all serve the same end: to import external structure that forces the analyst to confront the evidence they would rather ignore.

The Alarm TestDSI Advisory's named application of ACH and Popperian falsification turned reflexively on the security industry's own alarms, returning a calibrated SURVIVES, CALIBRATE or GENUINE-COUNTER verdict.: turning the method on ourselves

DSI Advisory built a named application of this discipline called the Alarm Test. The security industry runs on alarms: warnings that a technology, an adversary or a trend represents an urgent threat. Most such alarms are never tested; they are simply repeated. The Alarm Test takes a specific industry alarm and treats it as a hypothesis under ACH, then applies Popperian falsification, asking what evidence would have to appear for the alarm to be genuine, and what evidence already cuts against it. Crucially, it turns the method reflexively on the industry's own claims rather than only on external adversaries, which is the harder and less flattering direction.

A calibrated verdict, not a headline

The test yields one of three calibrated verdicts. An alarm SURVIVES when it withstands the attempt to falsify it. It is marked CALIBRATE when there is a real signal but the stated severity outruns the evidence. It is a GENUINE-COUNTER when the evidence actively contradicts it. DSI is careful to attribute honestly here: the Alarm Test is not a new method but a disciplined application of Heuer and Popper to a target the industry rarely examines, its own noise. The contribution is the target and the reflexive self-test, not a claim to have invented falsification.

ACH scores hypotheses by diagnostic evidence that is inconsistent with them.
ACH scores hypotheses by diagnostic evidence that is inconsistent with them.

Check your understanding

Answer each from memory. Your results are saved in this browser and count toward your readiness — sign in (account panel above) to keep them across devices.

  1. In Analysis of Competing Hypotheses, which evidence is most valuable?

  2. What Popperian idea underlies ACH?

  3. What is distinctive about how DSI's Alarm Test applies this tradecraft?

13 min · Core

ICD 203 Analytic Standards

Good method still needs a standard of expression. Intelligence Community Directive 203 sets out how US intelligence analysis must be written: with calibrated confidence, explicit alternatives, transparent sourcing and strict non-partisanship. This lesson examines those standards and why DSI Advisory writes to this bar, given that its readership includes the US intelligence community and elected officials.

~4 min

By the end you can

  • Summarise the core analytic standards codified in ICD 203.
  • Distinguish a calibrated confidence band from a bare probability guess.
  • Explain why analysis of alternatives and non-partisanship are structural, not stylistic.
  • Explain why DSI writes public analysis to the ICD 203 bar.

A standard for how judgement is expressed

A sound analytic process can still be communicated badly, in ways that mislead the reader about how much to trust the conclusion. The US Intelligence Community addressed this with Intelligence Community Directive 203, which codifies the analytic standards every finished product must meet. It is not a method for reaching judgements but a discipline for stating them honestly. Its requirements are deliberately unglamorous: they exist to prevent an analyst from sounding more certain, more sourced or more objective than the underlying work supports.

Calibrated confidence, not false precision

The most visible requirement is a calibrated confidence band. ICD 203Intelligence Community Directive 203, the US analytic standards requiring calibrated confidence, analysis of alternatives, transparent sourcing and non-partisanship in finished intelligence products. asks the analyst to separate two things that everyday language blurs: how likely an event is, and how much confidence the analyst has in that estimate given the sources. A judgement is stated with an explicit likelihood expression, and the confidence in it, high, moderate or low, is stated separately and justified by the quality and consistency of the evidence. This matters because a plausible-sounding number resting on a single unreliable source is far weaker than the same number backed by several independent streams, and the reader deserves to know which they are being handed.

Alternatives and sourcing on the record

Two further standards are structural rather than cosmetic. Analysis of alternatives requires the product to set out the credible competing explanations and say why the leading judgement was preferred, which is ACH carried through to the page and a direct defence against confirmation bias. Transparent sourcing requires that the basis for a judgement, and the strengths and weaknesses of that basis, be made visible, so a reader can weigh the evidence rather than take the conclusion on faith. Together they make the reasoning auditable: a second analyst can retrace the path and see where they would differ.

Non-partisanship as a discipline

ICD 203 also demands that analysis be independent of political consideration. The judgement must follow from the evidence, not from what any patron wishes were true, and it must not advocate for a policy. This is not neutrality for its own sake; it is what makes the product usable by decision-makers of any persuasion, because they can trust it has not been bent to a conclusion in advance.

Why DSI writes to this bar

DSI Advisory adopts the ICD 203 standards for its public analysis, and the reason is concrete. Its readership includes members of the US intelligence community and elected officials, an audience trained to spot an unsourced assertion, an uncalibrated confidence claim or a partisan tell. Writing to this bar is how an independent firm earns the right to be read by that audience. Independence is the firm's advantage, and independence is worthless if the analysis is sloppy, so the standard is not borrowed prestige but a working constraint: sourced, confidence-banded, alternatives on the record, and structural rather than partisan in its explanations.

The standards make reasoning auditable so a reader can retrace and trust it.
The standards make reasoning auditable so a reader can retrace and trust it.

Check your understanding

Answer each from memory. Your results are saved in this browser and count toward your readiness — sign in (account panel above) to keep them across devices.

  1. What does a calibrated confidence band under ICD 203 require an analyst to separate?

  2. Why is analysis of alternatives a structural standard rather than a stylistic flourish?

  3. Why does DSI Advisory write its public analysis to the ICD 203 standard?

13 min · Core

Reference-Class Forecasting

Even a careful analyst tends to forecast from inside the specific case, which reliably produces optimism and error. Reference-class forecasting corrects this by starting from the outside: find the class of comparable past cases and take their base rate as the anchor. This lesson explains the outside view, the planning fallacy it cures, and the work of Kahneman and Bent Flyvbjerg.

~4 min

By the end you can

  • Distinguish the inside view from the outside view in forecasting.
  • Define the planning fallacy and explain why it is so persistent.
  • Describe how a reference class and its base rate produce a better anchor.
  • Apply reference-class thinking to a foresight or risk estimate.

Two ways to forecast a case

Faced with estimating how a project or an event will unfold, an analyst can take one of two stances. The inside view builds the forecast from the details of the specific case: this team, this plan, these particular circumstances. It feels rigorous because it engages with the facts in front of you. The outside view ignores those details at first and asks a different question: what happened in the broad class of cases like this one. Kahneman and Tversky drew the distinction and found, repeatedly, that the inside view is seductive and wrong, while the outside view is unglamorous and far more accurate.

The planning fallacy

The clearest symptom is the planning fallacy: the well-documented tendency to predict that a task will go better and finish sooner than nearly identical past tasks actually did. It persists because each planner believes their case is special, that the delays and overruns that afflicted everyone else will not apply here. The details of the specific plan crowd out the memory of how such plans usually end. Confidence in the inside view is not a sign that it is right; it is the very mechanism by which it goes wrong, because attention to the specific case suppresses the base rate.

Anchoring on the base rate

Reference-class forecasting is the cure Kahneman proposed and Bent Flyvbjerg turned into practice. The procedure is deliberate. First, identify the reference class: the set of past cases that genuinely resemble the one at hand, broadly enough to be numerous but narrowly enough to be comparable. Second, establish the base rate: what actually happened across that class, the distribution of outcomes, costs or durations. Third, anchor the forecast on that base rate and adjust only for genuinely distinctive features of the present case, cautiously. The outside view becomes the starting point and the inside view a modest correction, not the other way round.

Flyvbjerg and the evidence

Flyvbjerg's study of major infrastructure projects, from rail lines to bridges, gave the method its hardest evidence. He found that inside-view estimates were optimistic to the point of systematic bias, with cost overruns and shortfalls the rule rather than the exception, and that anchoring on the actual record of comparable projects produced dramatically better forecasts. The lesson generalises well beyond construction. Whenever you estimate how long an adversary will take to develop a capability, how a crisis is likely to resolve, or how a technology will diffuse, the discipline is the same: find the reference class, respect its base rate, and treat your sense that this case is different as a hypothesis to be checked rather than a fact.

Why it is hard to do

Reference-class forecasting is simple to describe and difficult to practise, because it demands setting aside the compelling narrative of the specific case in favour of a dull statistic about cases the analyst may find less interesting. It also requires the discipline to build the class honestly rather than gerrymandering it to include only flattering precedents. Done well, it is one of the most reliable corrections available to a forecaster, and it pairs naturally with the accountability the next lesson describes.

Anchor on the base rate of comparable cases, then adjust cautiously.
Anchor on the base rate of comparable cases, then adjust cautiously.

Check your understanding

Answer each from memory. Your results are saved in this browser and count toward your readiness — sign in (account panel above) to keep them across devices.

  1. What does the outside view do that the inside view does not?

  2. Why does the planning fallacy persist even among careful, experienced planners?

  3. In reference-class forecasting, what is the correct role of the specific case's distinctive features?

14 min · Core

The Predictions Scorecard

Method and standards mean little without accountability. A forecaster who never records what they predicted, and never checks how it turned out, cannot improve and cannot be trusted. This lesson examines the discipline of scoring forecasts, Brier scores and calibration curves, Tetlock's evidence that scoring makes forecasters better, and DSI's practice of publishing the test, not just the conclusion.

~4 min

By the end you can

  • Explain why recording and scoring forecasts is essential to foresight.
  • Describe what a Brier score measures and why it rewards calibration.
  • Distinguish calibration from resolution in forecast performance.
  • Explain DSI's practice of publishing the test, not only the conclusion.

The forecast nobody checks

Most public forecasting is never scored. A commentator predicts, the moment passes, and hindsight bias quietly rewrites memory so that whatever happened feels as though it was expected. Without a record, there is no feedback, and without feedback there is no learning: the forecaster cannot tell a lucky guess from a sound judgement, and neither can the reader. The first act of accountable foresight is therefore mundane and uncomfortable, to write the prediction down, with a probability and a deadline, before the outcome is known.

Scoring with the Brier scoreA measure of forecast accuracy equal to the average squared difference between the predicted probability and the actual outcome; lower is better, and it penalises both being wrong and being overconfident.

Once forecasts are recorded as probabilities, they can be scored. The Brier score, which Philip Tetlock used throughout his forecasting research, measures the average squared distance between what you predicted and what happened, coded as one or zero. Lower is better. Its virtue is that it punishes two different sins at once: being wrong, and being overconfident. A forecaster who says ninety per cent and is wrong is penalised heavily, while one who honestly says fifty-five per cent is barely penalised whichever way it falls. The Brier score therefore rewards saying what you actually believe with appropriate humility, rather than performing certainty.

CalibrationThe property of a forecaster whose stated probabilities match observed frequencies, so that events assigned a 70 per cent chance occur about 70 per cent of the time; drawn as a curve against the diagonal of perfect calibration. and resolution

Two qualities make a good forecaster, and they are worth separating. Calibration asks whether your probabilities mean what they say: across all the times you said seventy per cent, did the event happen about seventy per cent of the time? A well-calibrated forecaster's confidence tracks reality, and this can be drawn as a calibration curve against the diagonal of perfect calibration. Resolution asks whether you were willing to move away from the base rate toward decisive predictions when the evidence allowed it, rather than hedging everything to safe middle numbers. A useful forecaster needs both: calibration keeps them honest, resolution keeps them informative. Tetlock's long tournaments showed that keeping score, and studying one's own curve, measurably improves both over time.

Publishing the test, not just the conclusion

DSI extends this accountability into how it publishes. Its Predictions ScorecardDSI's accountability instrument that records the firm's forecasts and returns to grade them in public, embodying the method-family principle of publishing the test rather than only the conclusion. is an instrument that records the firm's forecasts and returns to grade them, in public, rather than quietly dropping the ones that aged badly. The wider principle, which runs through the firm's method-family, is to publish the test, not just the conclusion: to show the reasoning, the sourcing and the falsification attempt so a reader can judge the verdict rather than accept it. The Alarm TestDSI Advisory's named application of ACH and Popperian falsification turned reflexively on the security industry's own alarms, returning a calibrated SURVIVES, CALIBRATE or GENUINE-COUNTER verdict. embodies the same discipline for individual claims, publishing the SURVIVES, CALIBRATE or GENUINE-COUNTER working rather than only the headline. Transparency about method is what lets an independent firm be held to account by its own readers.

The method-family as a whole

These pieces form one connected practice. DSI describes its tradecraft as a method-family that runs from asking the right Questions, through mapping the Full Threat Surface, through the Alarm Test that examines specific claims, to the Scorecard that holds forecasts to account over time. Each stage borrows honestly from the canon in this module, Heuer's ACH, ICD 203Intelligence Community Directive 203, the US analytic standards requiring calibrated confidence, analysis of alternatives, transparent sourcing and non-partisanship in finished intelligence products.'s standards, Kahneman and Flyvbjerg's outside view, and Tetlock's scoring, and turns it into a repeatable discipline. The through-line is the same conviction that opened the module: unaided judgement is unreliable, so the answer is structure, calibration and accountability, applied without exception to one's own work.

Recording and grading forecasts is the feedback loop that turns luck into learning.
Recording and grading forecasts is the feedback loop that turns luck into learning.

Check your understanding

Answer each from memory. Your results are saved in this browser and count toward your readiness — sign in (account panel above) to keep them across devices.

  1. What does the Brier score reward, and why is that useful?

  2. What does calibration measure in a forecaster's record?

  3. What does DSI's practice of 'publishing the test, not just the conclusion' mean?

Flashcards

Recall-first review of the load-bearing facts.

0 reviewed · 8 left

Ready to test yourself?

12 graded questions with real explanations. You commit a confidence before each reveal — that is how you find what you only think you know.

Start practice quiz →
Judgement, Calibration and Tradecraft — Strategic Foresight | Contested Futures Academy · The Contested Futures Institute