Mastering EpistemologyGuide · Map · Audio فا
The Sleep of Reason Produces Monsters, plate 43 of Los Caprichos by Francisco Goya (1799)
Part IV · Knowledge in Society and in the Mind · Chapter 14

The Psychology of Reasoning: How Minds Actually Work

“The human understanding when it has once adopted an opinion ... draws all things else to support and agree with it.”

— Francis Bacon, Novum Organum (1620), I.46
37 min read 59 min audio
Francisco Goya, The Sleep of Reason Produces Monsters, plate 43 of Los Caprichos, 1799
On the cover

Francisco Goya, The Sleep of Reason Produces Monsters, plate 43 of Los Caprichos, 1799, etching and aquatint. When reason sleeps, owls and bats crowd in. The psychology of reasoning studies what fills the space when careful thought switches off.

Part IV · Knowledge in Society and in the Mind · Chapter 14
The Psychology of Reasoning: How Minds Actually Work
37 min read
Listen to this chapter59 min · narrated
In this chapter
  1. Why psychology matters to epistemology
  2. Dual-process theories
  3. Heuristics and biases
  4. Confirmation bias and myside bias
  5. Motivated reasoning
  6. Belief perseverance and biased assimilation
  7. Fluency and the illusory truth effect
  8. Overconfidence and the illusion of explanatory depth
  9. The Dunning-Kruger effect and its critics
  10. Identity-protective cognition
  11. Mindreading: how we attribute knowledge
  12. The argumentative theory of reasoning
  13. Ecological rationality
  14. Replication and the psychology of psychology
  15. Debiasing: what works
  16. The scout mindset
  17. Forecasting and superforecasters
  18. Check your understanding
  19. Further reading

“Beliefs are hypotheses to be tested, not treasures to be guarded.”

— Philip Tetlock and Dan Gardner, Superforecasting (2015)

Epistemology tells us how we ought to reason. Psychology tells us how we actually reason. The gap between the two is where most of our errors live. Since the 1960s, cognitive psychologists have documented systematic, predictable patterns in human judgment: shortcuts that usually serve us well and sometimes lead us badly astray, and motivational forces that bend our reasoning toward what we want to believe.

This chapter surveys the most important findings and, crucially, what can be done about them. It also practices what it preaches: psychology has had its own replication crisis, and several famous findings have not held up. They are flagged here. A critical thinker should be critical of the science of critical thinking too.

Why psychology matters to epistemology

W. V. O. Quine argued that epistemology should become “a chapter of psychology” (see Quine and naturalized epistemology). Few epistemologists go that far, but most agree that empirical facts about cognition matter, for three reasons:

  1. “Ought” implies “can.” Epistemic norms should be norms that human beings can actually follow. A theory of rationality that requires unlimited memory and computing power is not a guide for people.
  2. Reliability is an empirical matter. If justification depends on the reliability of belief-forming processes (see Reliabilism), then we need to know which processes are reliable, and that is a question for psychology.
  3. Knowing our weaknesses helps us correct them. This is the practical aim of this chapter.

A debate runs through this field about how to interpret the evidence of human error. Keith Stanovich (Who Is Rational?, 1999) identified three positions:

  • Meliorists (Kahneman, Tversky, Stanovich): people’s reasoning often falls short of normative standards, and can be improved.
  • Panglossians (L. Jonathan Cohen, “Can Human Irrationality Be Experimentally Demonstrated?,” 1981): since our normative standards are ultimately derived from human intuitions, people cannot be systematically irrational; apparent errors are performance slips, misunderstandings of the task, or experimenter error.
  • Apologists (Gigerenzer and the “ecological rationality” program): apparent errors often reflect heuristics that are well adapted to real environments, even if they fail on artificial laboratory tasks.

Each position contains truth. The sections below try to present the evidence in a way that respects all three.

Dual-process theories

Many psychologists distinguish two kinds of thinking, popularized by Daniel Kahneman (Thinking, Fast and Slow, 2011), using the labels System 1 and System 2 introduced by Keith Stanovich and Richard West (2000).

System 1 (Type 1) System 2 (Type 2)
Fast Slow
Automatic, effortless Deliberate, effortful
Intuitive, associative Rule-based, reflective
Operates in parallel Operates serially
Does not require working memory Requires working memory
Generates impressions and feelings Can endorse or override those impressions

System 1 recognizes faces, reads emotions, completes “bread and...”, drives a car on an empty road, and answers “2 + 2.” System 2 multiplies 17 × 24, checks a complex argument, fills in a tax form, and resists an impulse.

The Cognitive Reflection Test (Shane Frederick, 2005) measures the tendency to override an intuitive but wrong answer. Try it:

  1. A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost?
  2. If it takes 5 machines 5 minutes to make 5 widgets, how long would it take 100 machines to make 100 widgets?
  3. In a lake, there is a patch of lily pads. Every day, the patch doubles in size. If it takes 48 days for the patch to cover the entire lake, how long would it take for the patch to cover half of the lake?
Answers
  1. 5 cents (not 10 cents: if the ball were 10 cents, the bat would be $1.10, and the total $1.20).
  2. 5 minutes (each machine makes one widget in 5 minutes).
  3. 47 days (the day before it covers the whole lake, it covers half).

Frederick found that intuitive wrong answers were common even among students at highly selective universities. The intuitive answer comes to mind immediately and feels right; getting it right requires noticing that it needs checking.

Caveats. The two-system picture is a useful simplification, not a precise model of the brain. Critics (David Melnikoff and John Bargh, “The Mythical Number Two,” 2018) argue that the features attributed to each system don’t cluster as neatly as the theory claims. Jonathan Evans and Keith Stanovich (“Dual-Process Theories of Higher Cognition,” 2013) responded by narrowing the defining feature of Type 2 processing to its use of working memory and its capacity for hypothetical thinking. Also, System 1 is not the “bad” system: expert intuition is System 1, and it can be highly reliable (see Intuition, emotion, and other candidate sources). The problem arises when System 1 answers a question it is not equipped for and System 2 fails to notice.

Heuristics and biases

In their landmark paper “Judgment under Uncertainty: Heuristics and Biases” (Science, 1974), Amos Tversky and Daniel Kahneman proposed that people answer hard questions about probability by substituting easier ones, using heuristics (mental shortcuts). Heuristics are often effective, but they produce biases: systematic, predictable deviations from normative standards.

Availability

The availability heuristic: judging how frequent or probable something is by how easily examples come to mind.

  • Are there more English words that begin with the letter K, or that have K as their third letter? Most people say words beginning with K, because they are easier to recall. In fact, in a typical text, words with K in the third position are more common (Tversky and Kahneman, 1973).
  • People overestimate the frequency of dramatic, well-publicized causes of death (tornadoes, homicides, accidents) and underestimate common, undramatic ones (diabetes, stroke, asthma) (Sarah Lichtenstein and colleagues, 1978).
  • After the September 11, 2001 attacks, many Americans avoided flying and drove instead. Gerd Gigerenzer estimated that the resulting increase in road traffic led to roughly 1,500 additional road deaths in the following year, more than the number of passengers killed on the four hijacked planes.

Availability is shaped by media coverage, which selects for rarity and drama (“if it bleeds, it leads”), and by personal experience, which is a tiny, unrepresentative sample. See Base rate neglect.

Representativeness

The representativeness heuristic: judging the probability that something belongs to a category by how much it resembles the typical member of that category.

It produces:

  • Base rate neglect (the engineer/lawyer and cab problems; see Chapter 9).
  • The conjunction fallacy (Linda; see Chapter 9).
  • Misconceptions of chance: people judge the coin-flip sequence H-T-H-T-T-H to be more likely than H-H-H-T-T-T, although both are equally likely, because the first “looks random.”
  • The law of small numbers: expecting small samples to resemble the population (see Chapter 9).

Anchoring

Anchoring: estimates are pulled toward an initial value (the anchor), even an obviously irrelevant one.

In Tversky and Kahneman’s classic demonstration (1974), a wheel of fortune was rigged to stop at either 10 or 65. Participants were first asked whether the percentage of African countries in the United Nations was higher or lower than the number on the wheel, and then asked to estimate the percentage. The median estimates were 25% for those who saw 10, and 45% for those who saw 65.

Anchoring affects experts too:

  • In a study by Gregory Northcraft and Margaret Neale (1987), real estate agents’ appraisals of a house were influenced by the listing price they were shown, though most agents denied that the listing price had influenced them.
  • Birte Englich, Thomas Mussweiler, and Fritz Strack (“Playing Dice with Criminal Sentences,” 2006) had experienced German judges read a case file, then roll a pair of dice (secretly loaded to show a total of either 3 or 9) and consider whether the sentence should be higher or lower than that number of months. Judges who rolled 9 gave substantially longer sentences on average (about 8 months) than judges who rolled 3 (about 5 months).

Anchoring is exploited deliberately in negotiation (the first offer), retail (“was $200, now $99”), and fundraising (suggested donation amounts).

Framing effects

Framing effects occur when logically equivalent descriptions of the same options lead to different choices.

Tversky and Kahneman’s Asian disease problem (“The Framing of Decisions and the Psychology of Choice,” 1981): the United States is preparing for an unusual disease expected to kill 600 people. Two programs are proposed.

Positive (gain) frame:

  • Program A: 200 people will be saved. (72% chose this.)
  • Program B: a one-third probability that all 600 will be saved, and a two-thirds probability that no one will be saved. (28%.)

Negative (loss) frame, given to a different group:

  • Program C: 400 people will die. (22%.)
  • Program D: a one-third probability that nobody will die, and a two-thirds probability that all 600 will die. (78%.)

Programs A and C are identical, as are B and D. Yet preferences reversed. People are risk-averse for gains and risk-seeking for losses, a pattern explained by Kahneman and Tversky’s prospect theory (1979).

Framing affects professionals. In a study by Barbara McNeil and colleagues (1982), patients, students, and physicians chose between surgery and radiation therapy for lung cancer. Surgery was much more attractive when outcomes were described in terms of the probability of surviving rather than the equivalent probability of dying.

Attribute framing: ground beef described as “75% lean” is rated more favorably than beef described as “25% fat” (Irwin Levin and Gary Gaeth, 1988).

Other well-documented biases

Bias Description Example
Loss aversion Losses loom larger than equivalent gains (roughly twice as large, in many estimates). Refusing a 50/50 bet to win $110 or lose $100.
Status quo bias Preferring the current state of affairs. Employees stay with whatever default pension option they were assigned.
Sunk cost effect Continuing an endeavor because of what has already been invested, not future prospects. Sitting through a bad movie because you paid for the ticket. See Sunk cost reasoning.
Hindsight bias After learning an outcome, believing you “knew it all along” (Baruch Fischhoff, 1975). “Everyone could see the crash coming.”
Outcome bias Judging a decision by its outcome rather than by the quality of the reasoning at the time. Praising a reckless bet that happened to pay off.
Fundamental attribution error Over-attributing others’ behavior to character and under-attributing it to circumstances (Lee Ross, 1977). “He’s late because he’s lazy,” not “because of the traffic.”
Halo effect A positive impression in one area colors judgments in others (Edward Thorndike, 1920). Attractive people are judged more intelligent and trustworthy.
Planning fallacy Underestimating the time, cost, and risk of future actions (Kahneman and Tversky, 1979). Students predicted they would finish their theses in about 34 days on average; they took about 56 (Roger Buehler and colleagues, 1994). The Sydney Opera House was planned to open in 1963 at a cost of 7 million Australian dollars; it opened in 1973 at a cost of 102 million.
Bias blind spot Seeing bias in others more readily than in oneself (Emily Pronin, Daniel Lin, and Lee Ross, 2002). “Their side is biased; I just see the facts.”
Affect heuristic Judging risks and benefits by how one feels about something (Paul Slovic). People who like a technology judge it both less risky and more beneficial; people who dislike it judge the reverse, although risks and benefits are not in fact inversely related.
Scope insensitivity Willingness to pay to prevent harm barely changes with the size of the harm. In one study, people were willing to pay about the same to save 2,000, 20,000, or 200,000 migrating birds (William Desvousges and colleagues, 1992).
Negativity bias Negative information weighs more than positive information of equal size. One bad review outweighs several good ones.
Belief bias Judging an argument’s validity by whether its conclusion is believable. Accepting an invalid syllogism with a plausible conclusion (Jonathan Evans and colleagues, 1983). See Validity and soundness.
In-group bias Favoring members of one’s own group, even when groups are arbitrary. In Henri Tajfel’s “minimal group” experiments, people favored their own group even when groups were assigned on trivial grounds.

Confirmation bias and myside bias

Confirmation bias is the tendency to seek, interpret, favor, and recall information in ways that confirm one’s existing beliefs. Raymond Nickerson’s review (“Confirmation Bias: A Ubiquitous Phenomenon in Many Guises,” 1998) called it perhaps the best known and most widely accepted notion of inferential error. Francis Bacon described it in 1620 (see the quote at the top of this chapter).

The 2-4-6 task (Peter Wason, 1960). Participants were told that the sequence 2-4-6 fits a rule the experimenter had in mind, and were asked to discover the rule by proposing other sequences; the experimenter would say whether each fit. Most participants formed a hypothesis (“numbers increasing by 2”) and tested it only with sequences that fit it (8-10-12, 20-22-24). Each got a “yes,” and they announced their hypothesis confidently. It was wrong. The rule was simply “any three increasing numbers.” To discover this, they needed to test sequences their hypothesis predicted would fail (1-2-3, or 5-20-100).

The selection task (Wason, 1966). You are shown four cards. Each has a letter on one side and a number on the other. The visible sides show:

E K 4 7

Rule: “If a card has a vowel on one side, then it has an even number on the other side.” Which cards must you turn over to test whether the rule is true?

Answer

E and 7. E, because if it has an odd number, the rule is false. And 7, because if it has a vowel on the other side, the rule is false. Most people choose E and 4, but the 4 cannot refute the rule (the rule says nothing about what’s behind even numbers). In Wason’s studies, fewer than 10% of participants chose correctly. This is the logic of modus tollens: to test “if P then Q,” look for not-Q cases, and check that they are not P.

Interestingly, people do much better with a logically equivalent version about social rules. “If a person is drinking beer, they must be over 18.” Cards: beer, cola, 25 years old, 16 years old. Most people correctly check beer and 16 (Richard Griggs and James Cox, 1982). Leda Cosmides (1989) argued that humans have specialized reasoning for detecting cheaters on social contracts.

Joshua Klayman and Young-Won Ha (1987) pointed out that testing cases your hypothesis predicts will succeed (a positive test strategy) is not always irrational. It is often a good strategy, and it fails mainly when the true rule is broader than your hypothesis, as in the 2-4-6 task. The lesson is not “never seek confirmation” but “make sure you also look for disconfirmation.”

Myside bias

Myside bias is the tendency to evaluate evidence, generate evidence, and test hypotheses in a manner biased toward one’s own prior opinions and attitudes. Keith Stanovich (The Bias That Divides Us, 2021) makes a surprising observation: unlike most biases, myside bias is not substantially reduced by intelligence or education. Smart people are just as biased toward their own side; they are simply better at generating arguments for it.

David Perkins (1985) asked people of different educational levels to reason about controversial issues. More educated people produced more arguments, but the extra arguments were almost all on their own side. Education made people better advocates, not better judges.

Motivated reasoning

Motivated reasoning is reasoning driven by a desire to reach a particular conclusion, not by a desire to find the truth. Ziva Kunda (“The Case for Motivated Reasoning,” 1990) distinguished accuracy goals (wanting to be right) from directional goals (wanting to reach a particular conclusion). She argued that directional goals influence reasoning by biasing which memories, beliefs, and rules of inference we access, but only within limits: people are constrained by an illusion of objectivity. They need to be able to construct a justification that would persuade a dispassionate observer. So motivated reasoning is not simply believing what we want; it is building a case for what we want, and believing the case.

Asymmetric scrutiny. Thomas Gilovich (How We Know What Isn’t So, 1991) described the pattern memorably: for conclusions we like, we ask “Can I believe this?” and look for any supporting evidence. For conclusions we dislike, we ask “Must I believe this?” and look for any reason to doubt. Peter Ditto and David Lopez (1992) demonstrated this experimentally. Participants were told they were being tested for a (fictitious) enzyme deficiency linked to later pancreatic disorders, using a test strip that would change color. Participants told that a color change indicated a healthy result waited longer for the strip to change, and when it didn’t, were more likely to re-test it; those who received the unwelcome result rated the test as less accurate.

Cognitive dissonance. Leon Festinger (A Theory of Cognitive Dissonance, 1957) proposed that holding inconsistent beliefs, or acting against one’s beliefs, creates an unpleasant tension that people reduce by changing their beliefs, often in self-serving ways. His study of a doomsday group (When Prophecy Fails, 1956) found that when the prophesied flood failed to arrive, some members became more committed, reinterpreting the failure as a success (see Ad hoc hypotheses).

Motivated numeracy. Dan Kahan, Ellen Peters, Erica Dawson, and Paul Slovic (“Motivated Numeracy and Enlightened Self-Government,” 2017) gave participants a tricky data-interpretation problem. When it was presented as a study about a skin cream, participants with higher numeracy were more likely to answer correctly. When the same numbers were presented as a study about gun control, accuracy depended on whether the correct answer fit participants’ political views, and the most numerate participants showed the largest partisan gap. Some later replications have found weaker effects, but the finding fits the broader pattern: skill at reasoning can be put to work in the service of the conclusion one wants.

Belief perseverance and biased assimilation

Belief perseverance: beliefs often survive the discrediting of the evidence that created them.

Lee Ross, Mark Lepper, and Michael Hubbard (1975) had participants try to distinguish real suicide notes from fake ones, and gave them false feedback that they had done very well or very badly. Afterwards, participants were fully debriefed: the feedback had been assigned at random and meant nothing. Yet participants who had been told they did well continued to rate their ability as higher than those told they did badly. Once people have constructed explanations for a belief (“I’m good at this because I’m empathetic”), the explanations remain even after the original evidence is gone.

Biased assimilation: people interpret mixed evidence as supporting what they already believe. Charles Lord, Lee Ross, and Mark Lepper (1979) gave supporters and opponents of capital punishment descriptions of two studies, one apparently supporting the death penalty’s deterrent effect and one apparently undermining it. Each side rated the study that supported its view as more convincing and better conducted, and found flaws in the other. Participants also reported that their views had become more extreme. (Later research suggests that people’s reported change may overstate actual change in attitudes, and a large study by Andrew Guess and Alexander Coppock, 2020, found little evidence that balanced information generally causes backlash. The biased evaluation of evidence, however, is well replicated.)

The continued influence effect, in which retracted misinformation continues to shape people’s reasoning, is discussed in Chapter 12.

Fluency and the illusory truth effect

Processing fluency, the ease with which information is processed, affects judgments of truth, liking, and confidence.

  • Rhyme as reason. Matthew McGlone and Jessica Tofighbakhsh (2000) found that aphorisms that rhyme (“Woes unite foes”) were judged more accurate than non-rhyming versions with the same meaning (“Woes unite enemies”).
  • The illusory truth effect. Repeated statements are judged more likely to be true than new ones (Lynn Hasher, David Goldstein, and Thomas Toppino, 1977). The effect occurs even for statements that contradict what people know (Lisa Fazio and colleagues, 2015), and even a single prior exposure to a fake news headline increases its perceived accuracy (Gordon Pennycook, Tyrone Cannon, and David Rand, 2018). Repetition is the propagandist’s oldest tool.
  • The Moses illusion. Asked “How many animals of each kind did Moses take on the ark?”, most people answer “two,” failing to notice that it was Noah (Thomas Erickson and Mark Mattson, 1981). When a question fits our expectations fluently, we don’t check its presuppositions.

Overconfidence and the illusion of explanatory depth

Don Moore and Paul Healy (“The Trouble with Overconfidence,” 2008) distinguished three forms of overconfidence:

  1. Overestimation: thinking you are better than you are (at a test, a task, a skill).
  2. Overplacement: thinking you are better than others. In a study by Ola Svenson (1981), a large majority of American student drivers rated themselves as safer and more skilled than the median driver in their group, which is statistically impossible for a majority.
  3. Overprecision: excessive certainty that your beliefs are accurate. When people give “90% confidence intervals” for uncertain quantities, the true value falls outside their range far more than 10% of the time. Overprecision is the most robust and widespread form. Try the calibration exercise in Chapter 9.

The illusion of explanatory depth

Leonid Rozenblit and Frank Keil (2002) asked people to rate how well they understood how everyday devices work: zippers, flush toilets, cylinder locks, sewing machines. Then they asked them to write a detailed, step-by-step explanation of how each works. After trying, people’s ratings of their own understanding dropped sharply. We confuse familiarity with a thing, or knowing that experts understand it, with understanding it ourselves.

Philip Fernbach, Todd Rogers, Craig Fox, and Steven Sloman (“Political Extremism Is Supported by an Illusion of Understanding,” 2013) extended this to politics. They asked people about policies such as cap-and-trade and a flat tax. After being asked to explain how each policy would work, step by step, participants rated their own understanding lower, and their positions became more moderate. Being asked to list reasons for their positions did not have this effect.

Steven Sloman and Philip Fernbach (The Knowledge Illusion, 2017) argue that the illusion arises because we live in a community of knowledge: we confuse knowledge that exists in other people’s heads, or in books, with knowledge in our own heads. That is usually fine (see Why knowledge is social), but it misleads us about our own competence.

The Dunning-Kruger effect and its critics

The Dunning-Kruger effect is probably the most famous finding in this chapter, and a good case study in how scientific results get distorted.

The original finding. Justin Kruger and David Dunning (“Unskilled and Unaware of It,” 1999) tested people on logical reasoning, grammar, and humor, and asked them to estimate their performance. Participants in the bottom quarter, who on average scored around the 12th percentile, estimated that they had scored around the 62nd percentile. Kruger and Dunning proposed a metacognitive explanation: the skills needed to perform well are the same skills needed to recognize good performance, so the incompetent are doubly cursed.

The popular version says something much stronger: that “stupid people think they are smart,” or that confidence falls as competence rises (the popular “Mount Stupid” graph, which does not come from Dunning and Kruger’s data).

The critique. Several researchers argued that much of the pattern is a statistical artifact. Two well-established facts combine to produce it:

  1. People’s self-assessments are only weakly correlated with their actual performance (so self-estimates for everyone cluster near the middle).
  2. Most people rate themselves as somewhat above average (the better-than-average effect).

Together, these guarantee that low performers will overestimate themselves and high performers will slightly underestimate themselves, even if everyone’s self-knowledge were equally poor. Simulations with random data reproduce the classic graph (Joachim Krueger and Ross Mueller, 2002; Edward Nuhfer and colleagues, 2016 and 2017). Gilles Gignac and Marcin Zajenkowski (2020), using methods designed to avoid these artifacts, found little evidence for the effect in the domain they studied. Dunning and his colleagues have replied that some of the effect survives the corrections.

What survives: people are generally poor judges of their own competence, and tend to overrate it; low performers overestimate by the largest margin, in absolute terms. In the actual data, low performers still rate themselves lower than high performers do; they do not think they are experts.

Identity-protective cognition

Why do people with the same information disagree so sharply on factual questions that have become politically charged, such as climate change, gun control, vaccines, or nuclear power?

Dan Kahan and his colleagues in the Cultural Cognition Project propose identity-protective cognition: people process information in ways that protect their standing in groups whose identity has become tied to particular positions. Kahan (“Why We Are Poles Apart on Climate Change,” Nature, 2012) argued that this is, in a narrow sense, individually rational. Any single person’s belief about climate change has a negligible effect on climate policy, but holding the “wrong” belief in one’s community can have real social costs. So individuals have strong incentives to believe what their group believes, even though society as a whole suffers when everyone reasons this way. Kahan called this the tragedy of the risk-perception commons.

A striking finding (Kahan and colleagues, Nature Climate Change, 2012): Americans with higher science literacy and numeracy were more polarized on climate change risk, not less. Greater reasoning skill was used to defend group positions more effectively.

Expressive responding. Some partisan “beliefs” may not be sincere beliefs at all, but expressions of loyalty. When John Bullock, Alan Gerber, Seth Hill, and Gregory Huber (2015) paid survey respondents small amounts for correct answers to factual questions about politics, and for admitting when they didn’t know, the partisan gaps in their answers shrank by more than half. Some apparent disagreement about facts is partisan cheerleading.

The debate. Gordon Pennycook and David Rand argue that inattention and a lack of reflective thinking, rather than identity-motivated reasoning, explain much susceptibility to misinformation (see Chapter 12). Both mechanisms probably operate, with different strengths on different questions.

Remedies suggested by the research:

  • Decouple facts from identity. Present evidence in ways that don’t signal membership in an opposing group.
  • Trusted messengers. People accept information more readily from sources they see as sharing their values.
  • Self-affirmation. In studies by Geoffrey Cohen, Joshua Aronson, and Claude Steele (2000), people who first wrote about an important value unrelated to the issue (affirming their identity) were then more open to evidence that challenged their views on a contested issue.
  • Keep your identity small. The programmer and essayist Paul Graham (“Keep Your Identity Small,” 2009) observed that discussions of religion and politics tend to be unproductive because they engage people’s identities, and advised: “the more labels you have for yourself, the dumber they make you.” Holding beliefs as conclusions rather than as parts of who you are makes it easier to change them.

Mindreading: how we attribute knowledge

Every day we judge, instantly and without calculation, whether someone knows something or merely thinks it: whether the colleague knows about the meeting, whether the driver has seen the cyclist. Psychologists call the capacity behind such judgments mindreading (or “theory of mind”): attributing hidden mental states such as wanting, believing, knowing, and pretending to others, on the basis of subtle cues like gaze and facial expression. Jennifer Nagel (Knowledge: A Very Short Introduction, ch. 8) argues that epistemologists should study this capacity, because our intuitions about cases like Gettier’s are its products.

Knowledge is easier to represent than belief.

  • Chimpanzees keep track of whether a rival knows or doesn’t know where food is hidden (Brian Hare, Josep Call, and Michael Tomasello, 2001), but no non-human animal has convincingly passed a test requiring it to track another’s false belief.
  • Human children pass explicit false-belief tests only around age four or five. In a classic version (the “unexpected contents” task, developed by Josef Perner, Heinz Wimmer, and colleagues in the 1980s), a child is shown a candy box, guesses it contains candy, and discovers it contains pencils. Most three-year-olds then say that another child who hasn’t looked will know it contains pencils, and many even say that they themselves thought all along that it contained pencils. Most five-year-olds get both questions right. (Studies suggest that infants have some implicit sensitivity to false beliefs, but the scope and robustness of this capacity are debated.)
  • Children in very different societies, from large cities to hunter-gatherer communities such as the Baka of Cameroon (Jeremy Avis and Paul Harris, 1991), pass through the same stages.
  • Children learn and use “know” earlier and more often than “think.”

To represent a false belief, you must hold in mind two pictures of the world at once, how it is and how someone takes it to be, and suppress the first. Knowledge, which must match the world, needs only one. Knowledge-first epistemologists cite these findings in support of their view (see Knowledge first).

The brain has specialized equipment for it. Reading stories about what characters think, want, and know selectively activates brain regions including the right temporo-parietal junction (Rebecca Saxe and Nancy Kanwisher, 2003), and damaging or temporarily disrupting this region impairs judgments that depend on others’ beliefs.

It has natural limits.

  • Capacity. We can track only a few levels of nested mental states. “Davis thinks that Lee knows that Smith doesn’t want Jones to find out about the job” is four levels; studies suggest most adults begin to fail at around five.
  • Epistemic egocentrism (the curse of knowledge). It is hard to set aside what you know when judging someone who knows less. In experimental markets, traders with private information failed to discount it when predicting how less-informed traders would behave, even when it cost them money (Colin Camerer, George Loewenstein, and Martin Weber, 1989). Once we know the outcome of a decision, we find it hard to judge the decision by what the decision-maker could have known (hindsight bias). Unlike many biases, egocentrism resists warnings and financial incentives.

Why it matters for epistemology and critical thinking.

  • Nagel suggests that egocentrism may explain the intuitions behind skepticism and contextualism. Once we are thinking about disguised mules or tricky lighting, we evaluate the naive zoo visitor or shopper as though they were entertaining those possibilities and irresponsibly ignoring them, and so we withdraw knowledge from them (see Thought experiments and intuitions and Subject-sensitive invariantism). If so, some of those intuitions are illusions to be handled with care.
  • The curse of knowledge makes experts poor judges of what novices understand, a major obstacle to good teaching and clear writing. Test your explanations on real beginners.
  • When judging whether someone “should have known,” ask what they could actually have known at the time, not what you know now.

The argumentative theory of reasoning

If human reasoning is so biased, what is it for? Hugo Mercier and Dan Sperber (“Why Do Humans Reason? Arguments for an Argumentative Theory,” Behavioral and Brain Sciences, 2011; The Enigma of Reason, 2017) proposed a provocative answer: reasoning did not evolve primarily to help individuals find the truth on their own. It evolved to produce arguments that persuade others and to evaluate the arguments others present to us, in a social world where communication is valuable but people have reasons to deceive.

This theory explains puzzling features of reasoning:

  • Myside bias is a feature, not a bug, of argument production. When you produce arguments, you act like a lawyer, looking for reasons for your side. That is what persuasion requires.
  • Laziness in producing arguments. People are satisfied with weak arguments for their own views, because in a dialogue, the other person will point out the weaknesses, and they can then improve their argument.
  • Evaluation of others’ arguments is much better. People are fairly good at spotting weak arguments when someone else makes them.

A striking demonstration of the last two points (Emmanuel Trouche, Petter Johansson, Lars Hall, and Hugo Mercier, “The Selective Laziness of Reasoning,” 2016): participants solved reasoning problems and gave arguments for their answers. Then they were shown other people’s answers and arguments, and asked to evaluate them. Through a trick, some of the “other people’s” arguments were actually the participants’ own arguments. Many participants rejected their own invalid arguments when they believed someone else had made them.

Groups reason better. On the Wason selection task, fewer than 10% of individuals typically get the right answer. David Moshman and Molly Geil (1998) found that when participants worked in small groups and discussed the problem, about three-quarters of groups found the right answer. Groups did not just adopt the majority view; they were persuaded by whoever had the correct argument.

Ecological rationality

Gerd Gigerenzer and his colleagues at the Max Planck Institute (the ABC Research Group; Simple Heuristics That Make Us Smart, 1999) offer a counterpoint to the heuristics-and-biases tradition. They argue:

  1. Heuristics are not just sources of error. Simple rules of thumb are often ecologically rational: well adapted to the structure of the environments in which they are used.
  2. Less can be more. In uncertain environments, simple heuristics that ignore information can predict better than complex models, because complex models overfit noise in the data (see Goodman’s new riddle of induction).
  3. Presentation matters. Many apparent biases shrink when information is presented in the formats our minds evolved to handle, such as natural frequencies (see Natural frequencies).

Examples:

  • The recognition heuristic. If you recognize one of two objects and not the other, infer that the recognized one has the higher value on the relevant criterion. Daniel Goldstein and Gigerenzer (2002) found that American students were slightly more accurate at judging which of two German cities was larger than which of two American cities was larger, because for German cities they could use recognition, while for American cities they recognized almost all of them and had to rely on other, less valid knowledge.
  • The 1/N rule. In a study of investment strategies (Victor DeMiguel, Lorenzo Garlappi, and Raman Uppal, 2009), simply dividing money equally among N assets performed as well out of sample as, and often better than, a range of sophisticated portfolio-optimization models, because the models’ parameters could not be estimated accurately enough from available data.
  • The gaze heuristic. A baseball player catching a fly ball does not calculate its trajectory. Players (and dogs catching frisbees) run so as to keep the angle of gaze to the ball constant, which brings them to where it will land.

Synthesis. Both traditions are right about something. Heuristics work well when the environment matches the one they are adapted to, and fail when it doesn’t: in novel situations, when information is presented in unfamiliar formats, or when someone is deliberately exploiting the heuristic (advertisers use anchoring, propagandists use repetition, scammers use urgency). A good reasoner uses heuristics, and knows when not to trust them.

Replication and the psychology of psychology

Psychology was at the center of the replication crisis. Some influential findings that have failed large replication attempts or shrunk substantially:

  • Ego depletion (willpower as a limited resource): a large multi-lab replication in 2016 found an effect near zero.
  • Power posing: failed to replicate its hormonal effects; one of the original authors disavowed it.
  • Elderly priming (John Bargh and colleagues, 1996: participants primed with words related to old age walked more slowly afterward): a replication by Stéphane Doyen and colleagues (2012) found the effect only when the experimenters expected it.
  • The facial feedback effect in its classic pen-in-mouth form: a multi-lab replication (Eric-Jan Wagenmakers and colleagues, 2016) did not reproduce it, though later studies have debated whether a smaller effect exists.
  • The backfire effect in response to factual corrections: found to be rare in larger studies (Thomas Wood and Ethan Porter, 2019).
  • The marshmallow test: the association between a child’s ability to delay gratification and later achievement was much smaller in a larger, more diverse sample once family background was taken into account (Tyler Watts, Greg Duncan, and Haonan Quan, 2018).
  • The Stanford Prison Experiment (Philip Zimbardo, 1971), which was never a controlled experiment, has been criticized on the basis of archival evidence that the guards were coached toward harsh behavior (Thibault Le Texier, 2019).

Many classic findings have held up well. The Many Labs 1 project (Richard Klein and colleagues, 2014), which ran 13 classic effects across 36 laboratories, replicated 10 of them, including anchoring and the gain/loss framing effect. Findings such as the conjunction fallacy, base rate neglect, the Wason selection task results, overprecision, and the illusory truth effect are well established.

This chapter has tried to report the current state of evidence. But the state of evidence changes. Treat claims about cognitive biases, including those in this guide, with the same critical standards you would apply to any other scientific claims.

Debiasing: what works

Simply learning about biases does little to reduce them. One reason is the bias blind spot: people readily recognize biases in others while believing themselves to be relatively free of them. What does work, to some degree?

1. Consider the opposite. Charles Lord, Mark Lepper, and Elizabeth Preston (“Considering the Opposite,” 1984) found that instructing people to be “as objective and unbiased as possible” did little to reduce biased assimilation. But instructing them to ask themselves, at each step, “Would I have made the same evaluation had exactly the same study produced results on the other side of the issue?” substantially reduced it. Specific procedures beat general exhortations.

2. Generate alternatives. Before settling on an explanation, list at least two or three alternatives, including boring ones (chance, measurement error, selection). See Inference to the best explanation.

3. The premortem. Gary Klein (“Performing a Project Premortem,” Harvard Business Review, 2007) recommends that before a team commits to a plan, members imagine that it is a year later and the plan has failed, then write down all the reasons why. This “prospective hindsight” makes it easier to identify risks. Deborah Mitchell, J. Edward Russo, and Nancy Pennington (1989) found that imagining an event had already occurred increased people’s ability to identify reasons for it.

4. Take the outside view. Start with base rates for similar cases before considering the specifics (see The problem of the priors). Reference class forecasting, developed by Bent Flyvbjerg for infrastructure projects, applies this systematically.

5. Use better formats. Natural frequencies instead of conditional probabilities; absolute risks instead of relative ones; visual displays of uncertainty (Gigerenzer).

6. Training can work, modestly. Richard Nisbett and colleagues found that training in statistical reasoning (such as the law of large numbers) transferred to everyday problems (Geoffrey Fong, David Krantz, and Nisbett, 1986). Carey Morewedge and colleagues (2015) found that a single session with an instructional video or a serious computer game reduced several biases (including confirmation bias, anchoring, and the bias blind spot), with effects lasting at least two to three months, and the game was more effective than the video.

7. Checklists and procedures. Atul Gawande (The Checklist Manifesto, 2009) describes how simple checklists reduce errors in surgery and aviation. A study of the World Health Organization’s Surgical Safety Checklist in eight hospitals (Alex Haynes and colleagues, New England Journal of Medicine, 2009) found that deaths fell from 1.5% to 0.8% and complications from 11% to 7% after the checklist was introduced. Procedures work because they don’t rely on the person remembering to be unbiased at the moment of decision.

8. Accountability, of the right kind. Jennifer Lerner and Philip Tetlock (“Accounting for the Effects of Accountability,” 1999) found that being accountable to an audience whose views are unknown, before forming one’s opinion, promotes more careful, balanced thinking. Being accountable to an audience whose views are known promotes conformity to those views; accountability after a decision promotes self-justification.

9. Blind analysis. Physicists routinely conduct blind analyses, hiding the final result from themselves (for example, by adding a secret offset to the data) until all decisions about the analysis have been made, so that they cannot be tempted to tune the analysis toward a desired result. Robert MacCoun and Saul Perlmutter (“Hide Results to Seek the Truth,” Nature, 2015) argued that other sciences should adopt the practice. The underlying idea, decide your method before you know the answer, is useful in everyday life too.

10. Adversarial collaboration. Daniel Kahneman promoted a practice in which researchers who disagree design an experiment together, agreeing in advance what results would support each side (Barbara Mellers, Ralph Hertwig, and Kahneman, “Do Frequency Representations Eliminate Conjunction Effects?,” 2001).

11. Structure group deliberation. Collect independent judgments before discussion; assign a devil’s advocate; seek diverse members; let junior people speak first (see The wisdom and madness of crowds).

12. Slow down on things that matter. Simple prompts to consider accuracy reduce the sharing of misinformation. And explaining mechanisms, not just reasons, reduces overconfidence and extremity.

The scout mindset

Julia Galef (The Scout Mindset, 2021) offers a useful image. Much of our reasoning is in soldier mindset: its purpose is to defend our beliefs against threatening evidence, like a soldier defending a position. The alternative is scout mindset: the motivation to see things as they are, like a scout whose job is to map the terrain accurately, whatever it looks like.

Galef argues that soldier mindset has real emotional and social benefits (comfort, self-esteem, morale, belonging), but that we overestimate them and underestimate the benefits of seeing clearly. Signs that you have scout mindset (paraphrasing her list):

  1. Do you tell other people when you realize they were right?
  2. How do you react to personal criticism? Do you actually seek it out?
  3. Do you ever prove yourself wrong?
  4. Do you take precautions to avoid fooling yourself?
  5. Do you have any good critics: people who disagree with you whom you consider reasonable?

Galef also proposes a set of thought experiments to catch motivated reasoning in the act:

Test Question to ask yourself
The double standard test Am I judging this person or group by a different standard than I’d use for another person or group?
The outsider test Imagine someone else stepped into my situation. What would I expect them to do?
The conformity test If the people I respect no longer held this view, would I still hold it?
The selective skeptic test If this evidence supported the other side, how credible would I find it?
The status quo bias test If my current situation were not the status quo, would I actively choose it?

Forecasting and superforecasters

Philip Tetlock’s work on expert political judgment (see Experts and novices) showed that most expert forecasters did little better than chance on long-range political questions. His later work showed that some people can forecast remarkably well, and that the skill can be learned.

From 2011 to 2015, the US intelligence community’s research agency (IARPA) ran a forecasting tournament. Participants forecast hundreds of real geopolitical questions (“Will North Korea conduct a nuclear test before the end of the year?”), and their accuracy was measured with Brier scores (see Calibration and scoring rules). Tetlock and Barbara Mellers’s Good Judgment Project, which recruited thousands of volunteers, won by a wide margin. Its best performers, the top 2% whom Tetlock called superforecasters, were reported to outperform professional intelligence analysts with access to classified information.

What made superforecasters good (Tetlock and Gardner, Superforecasting, 2015)?

  • Actively open-minded thinking (a concept from the psychologist Jonathan Baron): treating beliefs as hypotheses to be tested, and seeking out evidence against them. This trait predicted accuracy.
  • Numeracy and granularity: distinguishing between 60% and 65% probability, not just “likely.”
  • Frequent, small updates: revising forecasts often, in small increments, as new information arrived, neither overreacting nor underreacting.
  • Fermi-izing: breaking questions into smaller, more tractable parts (named after the physicist Enrico Fermi, who could estimate quantities like “the number of piano tuners in Chicago” by multiplying reasonable estimates of the components).
  • Outside view first, then inside view: starting from base rates, then adjusting for the specifics.
  • Synthesis of perspectives: combining many viewpoints, which Tetlock called “dragonfly eye.”
  • Teaming: working in teams that shared information and challenged one another improved accuracy.
  • Training: a short tutorial in probabilistic reasoning and common biases improved accuracy by around 10% over the course of a year.

Tetlock and Gardner’s “Ten Commandments for Aspiring Superforecasters” (paraphrased) is a good summary of practical epistemology:

  1. Triage: focus on questions where effort will pay off.
  2. Break seemingly intractable problems into tractable sub-problems.
  3. Strike the right balance between inside and outside views.
  4. Strike the right balance between under- and overreacting to evidence.
  5. Look for the clashing causal forces at work in each problem.
  6. Strive to distinguish as many degrees of doubt as the problem permits, but no more.
  7. Strike the right balance between under- and overconfidence, between prudence and decisiveness.
  8. Look for the errors behind your mistakes, but beware of rearview-mirror hindsight biases.
  9. Bring out the best in others and let others bring out the best in you.
  10. Master the error-balancing bicycle: these skills are learned by practice with feedback, not by reading.
  11. Don’t treat commandments as commandments.

Check your understanding

1Explain why the intuitive answer to the bat-and-ball problem is wrong, and what the problem measures.

Answer

If the ball cost 10 cents, the bat would cost $1.10 and the total $1.20. The correct answer is 5 cents (ball) and $1.05 (bat). The problem measures cognitive reflection: the tendency to notice that an immediate, intuitive answer needs checking, and to check it.

2Two groups receive the same information about a medical treatment, one framed in survival rates and one in mortality rates, and choose differently. What does this show, and how can you protect yourself?

Answer

A framing effect: logically equivalent descriptions produce different choices. Protect yourself by restating the options in the opposite frame, and in absolute numbers, and checking whether your preference changes.

3In the Wason selection task (E, K, 4, 7; rule: “if vowel, then even”), why is turning over the 4 useless?

Answer

The rule says nothing about what must be behind an even number. Whatever is behind the 4 (vowel or consonant) is consistent with the rule. Only a vowel with an odd number could falsify it, so you must check E (for an odd number) and 7 (for a vowel).

4What does Stanovich mean by saying myside bias is not reduced by intelligence? What protects against it?

Answer

People of higher cognitive ability show about as much myside bias as others; they are just better at producing arguments for their own side (as Perkins found with education). What helps is not ability but dispositions and methods: actively open-minded thinking, deliberately seeking the strongest opposing arguments, considering the opposite, and exposing one’s reasoning to critics.

5What is the “illusion of explanatory depth,” and how can it be used in a heated discussion?

Answer

People believe they understand how things work far better than they do, until they try to explain the mechanism step by step. In a discussion, asking each person (including yourself) to explain how a policy would produce its claimed effects often reveals gaps in understanding and tends to moderate extreme positions, whereas asking for reasons does not.

6Summarize the critique of the Dunning-Kruger effect.

Answer

Much of the classic pattern follows statistically from two facts: self-assessments correlate only weakly with actual performance, and most people rate themselves above average. Together these guarantee that low performers overestimate and high performers underestimate, even with no special metacognitive deficit among the unskilled. What survives is that people are generally poor judges of their own competence. The popular claim that incompetent people think they’re experts is not supported.

7According to the argumentative theory, why do groups outperform individuals on the Wason selection task?

Answer

Because reasoning is better at evaluating others’ arguments than at producing unbiased arguments of one’s own. In a group, a correct argument, once someone produces it, is recognized as correct by others, and weak arguments are criticized. Individuals, reasoning alone, tend to accept their own first intuitions and weak justifications.

8Why is “be unbiased” a poor debiasing instruction, and what works better?

Answer

Because people don’t perceive their own bias (the bias blind spot), so they believe they are already unbiased. Specific procedures work better: considering what you would think if the same evidence pointed the other way, generating alternative explanations, premortems, base rates, checklists, blind analysis, and structured group deliberation.

Further reading

Classics

  • Daniel Kahneman, Thinking, Fast and Slow (Farrar, Straus and Giroux, 2011). Some findings described in it, especially the priming studies in chapter 4, have not replicated, so read those parts with caution.
  • Amos Tversky and Daniel Kahneman, “Judgment under Uncertainty: Heuristics and Biases,” Science 185 (1974). Short and readable.
  • Thomas Gilovich, How We Know What Isn’t So (Free Press, 1991).
  • Richard Nisbett and Lee Ross, Human Inference (Prentice-Hall, 1980).

Recent books

  • Keith Stanovich, The Bias That Divides Us: The Science and Politics of Myside Thinking (MIT Press, 2021).
  • Hugo Mercier and Dan Sperber, The Enigma of Reason (Harvard University Press, 2017).
  • Steven Sloman and Philip Fernbach, The Knowledge Illusion (Riverhead, 2017).
  • Julia Galef, The Scout Mindset (Portfolio, 2021).
  • Steven Pinker, Rationality: What It Is, Why It Seems Scarce, Why It Matters (Viking, 2021).
  • Philip Tetlock and Dan Gardner, Superforecasting: The Art and Science of Prediction (Crown, 2015).
  • Gerd Gigerenzer, Gut Feelings (Viking, 2007) and Risk Savvy (Viking, 2014).
  • Daniel Kahneman, Olivier Sibony, and Cass Sunstein, Noise: A Flaw in Human Judgment (Little, Brown, 2021). On random variability in professional judgment, a companion problem to bias.

On replication

  • Stuart Ritchie, Science Fictions (Metropolitan, 2020).
Go deeper

Concepts from this chapter

Each has its own page with the key idea, objections and replies, common mistakes, and a self-check, in English and Persian.