Mental Models

Why the Models You Learned Stop Working: What Degrades in a Thinking Toolkit

You can define twenty mental models and still make exactly the decision you would have made without any of them. That is the normal outcome, not a personal failing, and it is the reason so many people who read seriously about thinking cannot point to a single decision that went differently because of it.

The usual diagnosis is that you have forgotten the material, and the usual prescription is to review it. Both are wrong. The definitions are the durable part of a thinking toolkit — they barely decay at all. What degrades is retrieval, calibration, and fit, and only one of the three is repaired by reading anything. Knowing which one has slipped tells you what maintenance is worth doing, and exposes how much of the standard advice maintains the only part that was never in trouble.

What actually degrades, and what doesn't

Test yourself on definitions after a year away and you will do better than you expect. You will still produce a serviceable account of the sunk-cost fallacy, still recognise inversion when you see it named, still explain what a base rate is. Concepts with a clean structure and a memorable name are unusually sticky.

That stickiness is exactly what makes self-assessment useless here. The check you can run at your desk — do I still know what this means? — measures the component that was never degrading. Meanwhile three other things quietly slip, and none of them show up on that test.

1. Retrieval: the model is in your notes, not in your hand

Retrieval failure is knowing something perfectly well and not producing it at the moment it would have helped. It is the gap between recognition and recall, and it is the single largest reason a well-read thinker decides like an unread one.

Consider a manager choosing between two final candidates. She has read about base rates. Two weeks after the hire goes badly she can explain exactly that she should have asked how often people with that profile succeed in that role, instead of reasoning from one unusually good interview. She was not short a model. She was short a cue — nothing in the situation said this is a base-rate problem. The habit itself is covered in our guide to thinking in base rates; having read it is not the same as having it arrive on time.

The limit worth naming: retrieval is not fixed by knowing the model better. Someone who can define anchoring in more detail is not measurably more likely to notice an anchor being set on them. Depth of understanding and speed of retrieval are close to independent, which is why studying harder transfers so poorly.

2. Calibration: your sense of how often you are right

Calibration is the match between your confidence and your accuracy. It has the fastest and quietest decay of the three, because the feedback that would maintain it mostly never arrives.

Outcomes come back late, noisily, and incompletely. You find out how the hire worked out; you never find out how the other candidate would have. Memory then does what memory does — confirmation bias favours evidence that you were right, and hindsight bias rewrites what you expected into something closer to what happened. Confidence keeps rising while accuracy stays flat, with no internal signal that anything has drifted.

Nothing you read repairs this. Calibration is a measurement problem, and measurement requires a record made before the outcome — which is why the decision journal keeps being recommended by people who agree on almost nothing else.

3. Fit: the model was right for a situation that has changed

Every model compresses, and the compression is tuned to an environment. When the environment moves, the model does not announce it.

The clearest case is any rule of thumb built from a personal reference class. "Projects like this take about six weeks here" is a base rate, and a good one — until the team, the tooling, or the client changes, at which point it is an out-of-date base rate wearing the authority of experience. The same happens to models about your field, your market, or how a particular kind of person behaves. They were earned, they were correct, and they expire without notice.

The failure mode is specific: a model that no longer fits does not feel uncertain. It feels like judgement. That is what makes fit the hardest of the three to audit, and why the repair is a periodic check of the reference class rather than a re-reading of the model.

The maintenance that is theatre

Each habit below feels like upkeep, is easy to measure, and produces the pleasant sensation of progress. Each one maintains definitions — the part that was not decaying.

  • Collecting more models. The most common substitute for practice. A latticework of mental models is genuinely the goal, but adding the fiftieth model while reaching for none of them is accumulation, not maintenance. As a rough guide, five or six models you actually pick up beat fifty you can define; the number is a rule of thumb, and the real test is whether you can name a decision each one changed.
  • Re-reading notes and highlights. Rereading produces fluency — the material feels familiar and easy, and the mind reads ease as mastery. It reliably strengthens the sensation of knowing without strengthening retrieval. Our guide to learning how to learn covers why testing yourself beats reviewing.
  • Naming the bias afterwards. Labelling a decision as confirmation bias once the outcome is in is a description, not a practice. Almost any bad outcome can be given a bias name in hindsight, which is exactly why doing so teaches nothing.
  • Spotting biases in other people. Enjoyable, socially rewarded, and the one form of practice that plausibly makes you worse — it trains the concepts as instruments of critique rather than as prompts for your own reasoning.
  • The annual review of your thinking. Too infrequent to correct anything and too far from the decisions to remember them honestly. By the time you sit down, the reasoning you meant to examine has been overwritten by the outcome.

The maintenance that works

Four habits, each aimed at something that genuinely degrades. None takes long, and all of them happen near decisions rather than near books.

Attach models to recurring decisions. This is the entire fix for retrieval. Do not try to remember the model — build the cue into the situation. A hiring template with one line at the top ("what is the base rate for this profile in this role?"). A purchase note with an opportunity-cost prompt. A kickoff that always opens with a pre-mortem. Cues attached to decisions you actually repeat beat any amount of review, because the model arrives when you need it rather than when you are studying.

Write three lines before the outcome. What I expect, why, and what would change my mind. That is the whole decision journal, and it is the only thing that maintains calibration. It has to be written before you know, and it has to record the reasoning rather than the choice — a record of what you decided without why is nearly useless later.

Review on outcomes, not on a schedule. Open the note when reality arrives, not every quarter. The comparison that teaches anything is between what you wrote and what happened, and it exists only at the moment the outcome lands.

Keep a surprise log. One line whenever something genuinely surprised you. Surprise is the cheapest available signal that a model no longer fits its environment, and the only early warning you get for the fit problem. Repeated surprises clustered in one area mean the reference class has moved.

One smaller habit earns its place: practise on decisions you are not inside. Reasoning through a choice someone else faces gives you repetitions without the motivated reasoning that comes from having a stake — the cheapest deliberate practice available to a solitary thinker. The broader frame sits in our guide to making better decisions.

How much maintenance is enough

There is no defensible number here, so take this as a rule of thumb rather than a finding: a working practice costs a few minutes per consequential decision, plus an hour or so a quarter looking back at the ones that resolved. Past that, maintenance competes with the deciding it exists to improve.

Over-maintenance has its own failure mode. Auditing every choice produces paralysis and a permanent sense that nothing has been examined enough, which is a worse state than deciding briskly with five reliable models. Match the depth of the examination to the stakes and the reversibility — the same instruction second-order thinking gives for the decisions themselves.

When the toolkit is not the problem

Two honest limits, before you conclude your thinking has degraded.

First, outcomes are not verdicts. A good decision can produce a bad result and a bad decision can produce a good one; judging your reasoning by results alone is the error usually called resulting, and it will talk you out of sound models because a fair coin landed badly. This is what the pre-outcome note protects: it lets you grade the reasoning separately from the roll.

Second, plenty of decisions do not go wrong for cognitive reasons at all. Where the incentives reward a particular conclusion, no amount of maintenance changes the answer — that is a structural problem wearing a thinking problem's clothes, treated by changing who decides rather than by sharpening how they think. The version that happens after the fact is covered in why smart people defend bad decisions.

Frequently asked questions

Do you forget mental models if you don't use them? Rarely the definitions — well-structured named concepts are unusually durable. What you lose is the ability to retrieve the right one in the moment, and any reliable sense of how accurate your judgement actually is. Both come back through use, not review.

Is a decision journal actually worth keeping? It is the only maintenance habit that repairs calibration, because calibration needs a record written before the outcome is known. Keep it minimal — expectation, reasoning, and what would change your mind — and revisit an entry only when the result arrives. Elaborate journals get abandoned.

How many mental models should I actively use? Fewer than you would guess. As a rough guide, a handful you genuinely reach for outperforms a large catalogue you can define, because the binding constraint is retrieval, not inventory. Add a new one when you can name the recurring decision it will attach to.

How do I know a model has stopped fitting my situation? Surprise is the signal. Repeated surprise in one area — estimates that keep missing in the same direction, people who keep behaving unexpectedly — usually means the reference class behind your rule of thumb has changed, not that your reasoning has weakened.

Does reading about biases make me less biased? Knowing a bias by name helps you recognise it in a description far more than it helps you catch it in yourself. Recognition and correction are different skills, and the corrective work is procedural — cues, checklists, records — rather than informational. Our guide to cognitive biases covers which ones respond to which countermeasures.

Maintenance done honestly is cheap and unglamorous: cues where decisions happen, three lines before the outcome, a look back when the result arrives, and a note of what surprised you. For the toolkit those habits are meant to keep sharp — every model and bias defined, applied, and given its limits — explore the mental-models encyclopedia on Build Mind.

Comments are disabled for this article.