A base rate is the background frequency of something in the relevant population — how often the thing happens in general, before you look at the specifics of your particular case. The disease affects one in a thousand people. Most restaurants close within a few years. Roughly nine in ten of these projects ship late. Those numbers are base rates, and the single most reliable upgrade you can make to your predictions is to start with one instead of skipping straight to the vivid details of the case in front of you.
The trouble is that vivid details are exactly what the mind reaches for. Told that a quiet, bookish man is "either a librarian or a farmer", most people guess librarian — because he sounds like a librarian — and quietly ignore that there are vastly more farmers than male librarians, so the base rate alone makes farmer the better bet. That reflex has a name, a well-documented history, and a fix. This guide covers all three.
What a base rate is
A base rate is a prior probability: the rate at which an outcome occurs across a whole reference class, independent of any specific evidence about the individual case. If 2% of loan applicants in a portfolio default, 2% is the base rate for default. If 70% of software features take longer than their first estimate, 70% is the base rate for overrun.
Base rates are the statistical background. Case-specific evidence — this applicant has a steady job, this feature looks simple — is the foreground. Good probabilistic reasoning combines the two: start from the background rate, then adjust it up or down for what is genuinely distinctive about the case. The formal machinery for doing this correctly is Bayes' theorem, but you do not need the algebra to get the core discipline: the base rate is where the estimate should start, not a footnote you add after you have already made up your mind.
Why your mind skips the base rate
The systematic tendency to underweight base rates in favour of specific, individuating information is called base-rate neglect, and it was documented in a series of experiments by the psychologists Daniel Kahneman and Amos Tversky. Their explanation was the representativeness heuristic: we judge how likely something is by how much it resembles our mental prototype, rather than by how common it actually is.
The bookish man resembles the prototype of a librarian, so "librarian" feels right — and resemblance crowds the actual population counts out of view entirely. The heuristic is fast and often works, but it is blind to frequency by design. That is the recurring shape of a cognitive bias: a useful shortcut with a predictable failure mode, the pattern covered across our guide to cognitive biases.
Two forces make the neglect worse. First, individuating detail is causally interpretable — a story about this specific case feels explanatory in a way a bare percentage never does, so the story wins the tug-of-war for attention. Second, base rates are often boring, hard to look up, or slightly embarrassing to consult ("surely my startup isn't like the others"). The information that would most improve the forecast is precisely the information the mind finds least compelling.
A worked example: the test that feels certain and isn't
Consider a purely illustrative case with round numbers chosen to show the mechanism. Suppose a disease affects 1 in 1,000 people, and a test for it is "95% accurate" — it correctly flags 95% of people who have the disease and wrongly flags only 5% of people who don't. Your test comes back positive. What is the chance you actually have the disease?
The intuitive answer is "about 95%", and it is badly wrong. Walk it through with a concrete population of 100,000 people:
- About 100 actually have the disease. The test catches roughly 95 of them: 95 true positives.
- The other 99,900 are healthy. The test wrongly flags 5% of them: about 4,995 false positives.
So roughly 5,090 people test positive, but only 95 of them are sick. Your chance of being ill given a positive result is about 95 in 5,090 — under 2%, not 95%. The rare-disease base rate dominates the answer, and ignoring it inflates the perceived risk more than fortyfold. Nothing changed about the test's accuracy; what changed is that we started from the base rate instead of the headline number. This is why a positive result on a screen for a rare condition usually means "retest", not "certainty" — the arithmetic, not optimism.
The lesson generalizes far past medicine. Any time a signal is imperfect and the thing it detects is rare, most of the signals it produces will be false alarms — fraud flags, resume filters, security alerts, "promising" early-stage bets. Skip the base rate and you will act on noise.
The outside view: base rates for one-off decisions
Base rates sound like a tool for repeatable statistical events, but their most valuable use is on the decisions that feel unique. Kahneman drew the distinction as the inside view versus the outside view. The inside view builds a forecast from the specifics of your case — your plan, your team, your reasons this time is different. The outside view ignores the details at first and asks: how did this class of project turn out for everyone who tried something similar?
The inside view is where the planning fallacy lives — the robust tendency to predict that your own project will run faster and cheaper than comparable projects reliably do, because you can see your plan from the inside and cannot see everyone else's identical optimism. The corrective, developed into a formal method by the planning researcher Bent Flyvbjerg under the name reference-class forecasting, is disciplined outside-view thinking: assemble the class of similar past efforts, take their actual distribution of outcomes as your starting point, and only then adjust for what is truly different about yours.
"How long will this renovation take?" invites a fantasy. "How long did the last ten comparable renovations actually take?" anchors you to reality. The move is always the same: find the reference class, read its base rate, and treat your optimism as something that must earn its adjustments rather than something that gets to set the estimate.
How to actually use base rates
Four steps turn this from a nice idea into a habit.
- Name the reference class. Ask "what is this a case of?" and find the broadest honest grouping of similar situations. This is the hard, judgment-heavy step — most base-rate errors are really reference-class errors.
- Get the base rate before you look at the specifics. Write down the background frequency first, deliberately, while you can still be objective — before the compelling story about your particular case has captured you.
- Adjust from there, sparingly. Now bring in the case-specific evidence, and move off the base rate only as much as genuinely diagnostic information justifies. Strong specifics warrant a real adjustment; a good feeling does not.
- Write the prediction down. Recording the base rate and your adjustment converts a vague hunch into a testable forecast you can review later — the same decision-journal discipline that anchors our framework for making better decisions.
Where base rates mislead
Anti-hype is the house rule, so here are the failure modes. The wrong reference class ruins everything. Pick a class too broad ("all businesses") or too narrow ("businesses exactly like mine, of which there are three") and the base rate misleads more than it helps. Choosing the class is a real judgment call, not a lookup.
Base rates go stale. A frequency drawn from a world that has since changed — a new technology, a shifted market, a different regime — is a fact about the past, not a reliable prior for the present. Ask whether the process that generated the number still operates.
Some events genuinely have no useful reference class. Truly novel situations resist the outside view; forcing a base rate onto them manufactures false precision. And base rates describe the group, never the individual with certainty — a 2% default rate does not tell you which borrowers default. The base rate sets the honest starting odds; it does not close the case.
Used with those limits in mind, base-rate thinking remains one of the highest-leverage habits in probabilistic reasoning: cheap, widely applicable, and a direct antidote to the vividness that hijacks our forecasts.
FAQ
What is a base rate in simple terms?
A base rate is how often something happens in general — the background frequency across a whole group before you consider the specifics of one case. If one in five restaurants in an area closes each year, that 20% is the base rate for closure, and it is the right place to start any prediction about a particular restaurant.
What is base-rate neglect?
Base-rate neglect is the well-documented tendency to ignore background frequencies in favour of specific, vivid details about the individual case. Studied by Kahneman and Tversky, it stems from the representativeness heuristic: we judge probability by resemblance to a mental prototype rather than by how common something actually is, so a case that "fits the type" feels likely even when the type is rare.
What is the difference between the inside view and the outside view?
The inside view forecasts from the details of your specific plan; the outside view forecasts from how similar past cases actually turned out. The inside view feels more informed but reliably produces overoptimistic estimates — the planning fallacy — because it cannot see that everyone else's plan looked equally promising from the inside. The outside view corrects this by starting from the base rate of the reference class.
How do I find the right base rate?
Define the reference class carefully: the group of past situations most genuinely similar to yours, broad enough to have real data but narrow enough to be relevant. Then look for the actual frequency of the outcome in that class — historical records, published statistics, or your own logged results — and check the number is recent enough that the world it came from still resembles today's.
Start from the odds
Every concept here — base rates, the representativeness heuristic, the planning fallacy, reference-class forecasting — has a full entry on Build Mind, with its definition, a worked example, and the failure mode that tells you when to stop trusting it. Explore the probability and base-rate entries on the Build Mind encyclopedia and put a base rate on the front of your next real prediction.