In this chapter
- What makes science special?
- The demarcation problem
- Inductivism and its limits
- Popper and falsificationism
- The Duhem-Quine problem
- Ad hoc hypotheses
- Kuhn and paradigms
- Lakatos and research programmes
- Feyerabend and Laudan
- Inference to the best explanation
- Theory-ladenness of observation
- Underdetermination
- Scientific realism and anti-realism
- Causation and causal inference
- Values in science
- Scientific consensus
- The replication crisis
- Pseudoscience and how to spot it
- Check your understanding
- Further reading
“Our knowledge can only be finite, while our ignorance must necessarily be infinite.”
— Karl Popper, “On the Sources of Knowledge and of Ignorance” (1960), in Conjectures and Refutations
Science is humanity’s most successful method for gaining knowledge about the natural world. But what exactly is the method? Why does it work? How should a non-scientist decide what to believe when scientists disagree, when a new study contradicts an old one, or when someone claims that “the science isn’t settled”?
This chapter covers the philosophy of science most useful for critical thinking: what distinguishes science from pseudoscience, how theories are tested and how they change, how to reason about explanations and causes, how to read the hierarchy of evidence, and what the replication crisis does and doesn’t show.
What makes science special?
A common answer is “science proves things.” It doesn’t, at least not in the mathematician’s sense. Scientific conclusions are fallible and revisable. What makes science special is not certainty but a set of practices that detect and correct error better than any alternative:
- Testing claims against observation, especially observation designed to catch the claim out.
- Controls, randomization, and blinding to exclude alternative explanations and the investigator’s own biases.
- Quantification and precision, which make predictions risky and errors visible.
- Public methods and data, so that others can check.
- Replication by independent investigators.
- Organized criticism: peer review, competing research groups, conferences, and a reward system that credits people for refuting others.
The sociologist Robert K. Merton (“The Normative Structure of Science,” 1942) described four norms of science, often remembered as CUDOS: Communism (findings are shared property), Universalism (claims are judged by impersonal criteria, not by who makes them), Disinterestedness (scientists act for the benefit of the common enterprise), and Organized Skepticism (all claims are subjected to critical scrutiny). Real science often falls short of these norms, but they describe the ideal that makes it work.
The historian Naomi Oreskes (Why Trust Science?, 2019) argues that the trustworthiness of science comes not from any single method, but from its social character: claims are vetted by a diverse community of experts through sustained critical scrutiny. This ties the philosophy of science to social epistemology.
The demarcation problem
What separates science from pseudoscience (astrology, homeopathy, creationism presented as science) and from non-science (mathematics, ethics, art)? This is the demarcation problem.
Karl Popper proposed falsifiability as the criterion (below). Many philosophers now think no single criterion works. Larry Laudan (“The Demise of the Demarcation Problem,” 1983) argued that no set of necessary and sufficient conditions separates science from non-science. More recent work (Massimo Pigliucci and Maarten Boudry, eds., Philosophy of Pseudoscience, 2013) treats “science” as a family resemblance concept: there is no single defining feature, but there are clear cases on either side and a cluster of relevant features.
Paul Thagard (“Why Astrology Is a Pseudoscience,” 1978) offered a useful diagnosis: a theory is pseudoscientific if it has been less progressive than alternative theories over a long period, faces many unsolved problems, and its community shows little concern with developing the theory to solve those problems, confirming it, or comparing it with alternatives. Astrology has barely changed in two thousand years, while astronomy has been transformed. Warning signs of pseudoscience are listed at the end of this chapter.
Inductivism and its limits
The popular picture of science, often traced (somewhat unfairly) to Francis Bacon, is inductivism: scientists collect observations without prejudice, and theories emerge from the facts by generalization.
This picture has serious problems (see A. F. Chalmers, What Is This Thing Called Science?, 4th ed. 2013):
- Observation is selective. Karl Popper used to begin lectures by telling his audience: “Take pencil and paper; carefully observe, and write down what you have observed!” They would ask: observe what? Observation needs a question, an interest, a hypothesis. There is no such thing as observation without selection.
- Theories go beyond observation. Theories posit unobservable entities (atoms, genes, gravitational fields) that no number of observations could be generalized into.
- The problem of induction. Even the simplest generalization from observed cases goes beyond the evidence. See Hume’s problem of induction.
Popper and falsificationism
Karl Popper (Logik der Forschung, 1934; English, The Logic of Scientific Discovery, 1959; Conjectures and Refutations, 1963) turned the inductivist picture upside down.
The asymmetry. No number of observations can prove a universal law true: after seeing a thousand white swans, the next might be black. (Europeans believed all swans were white until Dutch explorers encountered black swans in Western Australia in 1697.) But a single genuine counterexample can prove a universal law false. Verification and falsification are asymmetric. The logic of falsification is modus tollens: if the theory is true, we will observe O; we do not observe O; so the theory is false.
The criterion. A theory is scientific if and only if it is falsifiable: it forbids certain possible observations. The more it forbids, the more it says about the world, and the better it is, provided it survives testing.
Bold conjectures. Science progresses not by piling up confirmations (which are easy to find for almost any theory) but by proposing bold conjectures that make risky predictions, and then trying hard to refute them. A theory that survives severe tests is corroborated, but never proven.
Popper’s contrast. Popper’s formative experience came in Vienna around 1919. Einstein’s general relativity predicted that starlight passing near the Sun would be bent by a specific amount. Arthur Eddington’s 1919 eclipse expedition tested the prediction, and it could easily have failed. By contrast, the psychological theories of Freud and Alfred Adler seemed able to explain any human behavior. Popper’s example: a man pushes a child into water intending to drown it; another man sacrifices his life to save a child. Freud’s theory explains the first as repression and the second as sublimation; Adler’s theory explains the first as a feeling of inferiority needing to prove itself by crime, and the second as the same feeling needing to prove itself by daring. Popper recalled describing a case to Adler, who analyzed it confidently “in the light of my thousandfold experience.” Popper replied: “And with this new case, I suppose, your experience has become thousand-and-one-fold.” He concluded that the apparent strength of such theories, that they were always confirmed, was actually their weakness.
Problems for falsificationism.
- The Duhem–Quine problem: theories are never tested alone (next section).
- Scientists rightly don’t abandon theories at the first anomaly. Every major theory has faced anomalies throughout its history. If scientists had followed naive falsificationism, Newton’s theory would have been rejected at birth.
- Probabilistic hypotheses (“this coin is fair”) are not strictly falsified by any finite observation.
- Existential claims (“there exist black holes”) cannot be falsified by any finite set of observations, yet they are scientific.
- The role of confirmation. Wesley Salmon argued that when we rely on a well-corroborated theory to build a bridge, we treat its past success as evidence of future success, which is induction. See Responses to Hume.
Popper himself was aware of most of these problems and developed a more sophisticated position. His lasting insights are that testability is a virtue, that risky predictions are the strongest evidence, and that criticism drives the growth of knowledge.
The Duhem-Quine problem
The French physicist Pierre Duhem (The Aim and Structure of Physical Theory, 1906) observed that “an experiment in physics can never condemn an isolated hypothesis but only a whole theoretical group.” To derive a prediction from a theory, you always need auxiliary hypotheses: about initial conditions, about the instruments, about the absence of interfering factors. When the prediction fails, logic tells you that something in the group is false, but not what. W. V. O. Quine extended this to all beliefs: our statements face the tribunal of experience “only as a corporate body” (see Quine and naturalized epistemology).
Neptune and Vulcan. In the 1840s, the orbit of Uranus deviated from what Newtonian mechanics predicted. Should astronomers have rejected Newton? Urbain Le Verrier (and independently John Couch Adams) instead questioned an auxiliary assumption, that there were no other significant masses nearby. Le Verrier calculated where an unseen planet would have to be to cause the deviation. In 1846, Johann Galle pointed his telescope at the predicted position and found Neptune, within about one degree. Newton was triumphantly vindicated.
Then Le Verrier tried the same move with another anomaly: the orbit of Mercury precessed slightly faster than Newtonian theory predicted. In 1859 he proposed a small planet, Vulcan, orbiting inside Mercury’s orbit. Astronomers searched for decades, and several claimed sightings. Vulcan does not exist. The anomaly was finally explained in 1915 by Einstein’s general relativity, which replaced Newtonian gravitation.
The same strategy (protect the theory by revising an auxiliary) succeeded brilliantly once and failed once. Logic alone could not tell which case was which.
A modern example. In 2011, the OPERA experiment reported neutrinos apparently traveling faster than light, which would contradict special relativity. Rather than abandon relativity, the physicists themselves suspected an error in the experimental setup and asked others to find it. In 2012, the anomaly was traced to a loose fiber-optic cable and a clock problem.
Ad hoc hypotheses
An ad hoc hypothesis is one introduced solely to save a theory from refutation, with no independent evidence or testable consequences of its own.
Compare two responses to a failed prediction:
- Legitimate: “There must be an unseen planet at this position, with this mass.” (Testable: look there.)
- Ad hoc: “The prediction failed because the skeptics’ negative energy disrupted the psychic’s powers.” (Untestable: no independent way to detect negative energy.)
More examples of ad hoc rescues:
- An astrologer whose prediction fails says the client’s birth time must have been recorded incorrectly.
- A doomsday prophet whose predicted date passes announces that the faithful’s prayers postponed the end. (Leon Festinger, Henry Riecken, and Stanley Schachter studied a group that did this in When Prophecy Fails, 1956. After the predicted flood did not come, many members became more committed and began seeking converts. See Motivated reasoning.)
- Philip Henry Gosse’s “Omphalos” hypothesis (1857): God created the world recently, complete with fossils and geological strata that make it look ancient.
Criteria for telling a legitimate modification from an ad hoc one:
- Independent testability: does the modification have consequences beyond the anomaly it was introduced to explain?
- Novel predictions: does it predict something new that is then confirmed?
- Content: does it make the theory say more, or less? Ad hoc rescues usually drain content (“the effect only appears when no skeptic is watching”).
- Track record: does the theory need a new rescue after every test?
In Bayesian terms, an ad hoc auxiliary usually has a low prior probability and does nothing to raise it, so a theory that needs many of them becomes steadily less probable. Imre Lakatos made these ideas central to his account of science.
Kuhn and paradigms
Thomas Kuhn’s The Structure of Scientific Revolutions (1962) is one of the most cited academic books of the 20th century. Kuhn, a physicist turned historian, argued that the real history of science looked nothing like the textbook picture of steady accumulation or Popper’s picture of constant attempted refutation.
Paradigms. A mature science is organized around a paradigm: a set of shared achievements (exemplars, like Newton’s Principia or Lavoisier’s chemistry) that define the legitimate problems and methods of a field. In the 1969 postscript, Kuhn distinguished two senses: the disciplinary matrix (shared laws, models, values, and techniques) and exemplars (concrete problem-solutions students learn from).
Normal science. Most of the time, scientists do normal science: “puzzle-solving” within the paradigm, extending it and fitting nature to it. They do not try to refute the paradigm. When a result fails to fit, it is usually treated as a puzzle to be solved, or a failure of the scientist, not a refutation of the paradigm.
Anomaly and crisis. Sometimes anomalies accumulate that resist solution. Confidence in the paradigm erodes; scientists start proposing radical alternatives. This is a crisis.
Revolution. A new paradigm emerges and eventually wins acceptance. Examples: the Copernican revolution (Earth-centered to Sun-centered astronomy), the chemical revolution (Lavoisier’s oxygen theory replacing phlogiston), the Darwinian revolution, the Einsteinian revolution, and, in the 1960s, plate tectonics, which vindicated Alfred Wegener’s rejected 1912 hypothesis of continental drift after evidence from seafloor spreading and magnetic striping of the ocean floor accumulated.
Incommensurability. Kuhn argued that competing paradigms are partly incommensurable: they use key terms in different senses (“mass” means something different for Newton and Einstein), they recognize different problems as important, and they apply different standards. So there is no neutral algorithm for choosing between them. “The proponents of competing paradigms practice their trades in different worlds.” Joseph Priestley, who first isolated oxygen, never accepted Lavoisier’s interpretation and died defending phlogiston.
Theory choice without an algorithm. In “Objectivity, Value Judgment, and Theory Choice” (1977), Kuhn listed five values all scientists use to evaluate theories: accuracy, consistency, scope, simplicity, and fruitfulness. These are shared, but they can conflict, and scientists weigh them differently. So theory choice is rational but not mechanical, and reasonable scientists can disagree for a time.
Max Planck (Scientific Autobiography, 1950) offered a darker view: “A new scientific truth does not triumph by convincing its opponents and making them see the light, but rather because its opponents eventually die, and a new generation grows up that is familiar with it.” There is some empirical support for a mild version of this. Pierre Azoulay, Christian Fons-Rosen, and Joshua Graff Zivin (“Does Science Advance One Funeral at a Time?,” 2019) found that after the premature death of a star life scientist, publications by outsiders to the star’s field increase, and the new work is disproportionately well cited.
Was Kuhn a relativist? Many readers took Kuhn to show that science is irrational, driven by sociology rather than evidence. Kuhn denied this. He held that later theories are better puzzle-solvers than earlier ones and that science does make progress, though not necessarily toward “the truth” in a metaphysical sense.
Lakatos and research programmes
Imre Lakatos (“Falsification and the Methodology of Scientific Research Programmes,” 1970) tried to combine Popper’s rationalism with Kuhn’s history.
A research programme consists of:
- A hard core of central assumptions that the programme’s supporters decide not to question (for Newtonian mechanics, the three laws of motion and the law of gravitation). This is protected by a negative heuristic: don’t aim criticism at the hard core.
- A protective belt of auxiliary hypotheses, which absorbs the blows of anomalies and is modified as needed.
- A positive heuristic: a plan for developing the programme, telling researchers which problems to tackle and how to build more sophisticated models.
A research programme is:
- Progressive if its modifications lead to novel predictions, some of which are confirmed. The Newtonian prediction of Neptune is the classic example.
- Degenerating if its modifications are merely ad hoc patches that accommodate known facts without predicting anything new, or if its predictions keep failing and the theory lags behind the facts.
It is rational to work on progressive programmes and to abandon degenerating ones, though Lakatos admitted there is no precise time limit, since a degenerating programme can sometimes recover (atomism was out of favor for long periods). Paul Feyerabend pressed this point against him.
Lakatos’s framework is useful beyond science. Any belief system can be assessed by asking whether its responses to challenges are progressive or degenerating. Does it make new, testable claims that turn out to be true? Or does it only ever explain away the evidence after the fact?
Feyerabend and Laudan
Paul Feyerabend (Against Method, 1975) argued, from historical case studies such as Galileo’s defense of Copernicanism, that every proposed methodological rule has been broken by great scientists, often with good results. His provocative slogan was that the only principle that does not inhibit progress is “anything goes.” He later explained that this was “the terrified exclamation of a rationalist who takes a closer look at history.” He defended methodological pluralism and criticized what he saw as scientific chauvinism. Feyerabend is often misread as simply anti-science. A better reading: there is no fixed, algorithmic scientific method, and good science requires judgment, creativity, and tolerance for unorthodox ideas.
Larry Laudan (Progress and Its Problems, 1977) proposed that science is best understood as organized into research traditions and that progress should be measured by problem-solving effectiveness: how many empirical and conceptual problems a tradition solves, compared with its rivals. Laudan is also the source of the pessimistic meta-induction (see Scientific realism and anti-realism).
Bayesian philosophy of science offers a unifying framework. Many of the insights above can be expressed in Bayesian terms (see Chapter 9): risky predictions give large likelihood ratios (Popper); failed predictions spread their blame across the hypotheses and auxiliaries in proportion to their prior probabilities (Duhem–Quine, as analyzed by Jon Dorling in 1979 and Michael Strevens in 2001); ad hoc hypotheses have low priors (Lakatos); and scientists with different priors can reasonably disagree until evidence accumulates (Kuhn).
Inference to the best explanation
Inference to the best explanation (IBE), or abduction, is inferring that a hypothesis is true because it provides the best explanation of the available evidence (see Deduction, induction, and abduction). Charles Sanders Peirce described abduction; Gilbert Harman named IBE (“The Inference to the Best Explanation,” 1965); Peter Lipton developed the most thorough account (Inference to the Best Explanation, 1991; 2nd ed. 2004).
IBE is everywhere: in science, medicine (diagnosis), law (which story best explains the evidence?), history, detective work, car repair, and everyday life (the best explanation of the wet street is that it rained).
What makes an explanation “best”? The standard explanatory virtues:
| Virtue | Question to ask | Example |
|---|---|---|
| Scope | How much of the evidence does it explain? | Natural selection explains biogeography, fossils, anatomy, embryology, vestigial organs. |
| Precision | Does it explain the details, quantitatively? | General relativity predicts Mercury’s precession to within observational error. |
| Simplicity (parsimony) | Does it avoid unnecessary assumptions and entities? | One common cause rather than many coincidences. See Ockham’s razor. |
| Unification | Does it bring apparently different phenomena under one account? | Newton unified falling apples and orbiting planets. |
| Mechanism | Does it specify how the cause produces the effect? | Germ theory identifies the microorganisms that cause disease. |
| Fit with background knowledge | Is it consistent with what we already know well? | A diagnosis consistent with the patient’s history and physiology. |
| Fertility | Does it suggest new predictions and research? | Plate tectonics predicted patterns of earthquakes and seafloor magnetism. |
| Few ad hoc assumptions | Does it need special pleading to handle awkward facts? | See Ad hoc hypotheses. |
Lipton distinguished the likeliest explanation (the most probable given the evidence) from the loveliest (the one that would provide the most understanding). IBE claims that loveliness is a guide to likeliness.
Darwin’s IBE. Darwin explicitly argued this way in the Origin of Species (6th ed., 1872): “It can hardly be supposed that a false theory would explain, in so satisfactory a manner as does the theory of natural selection, the several large classes of facts above specified.”
Semmelweis and childbed fever. A classic case (analyzed by Carl Hempel in Philosophy of Natural Science, 1966). In the 1840s at the Vienna General Hospital, women in the First Maternity Division, staffed by doctors and medical students, died of puerperal (“childbed”) fever at a far higher rate than women in the Second Division, staffed by midwives. Ignaz Semmelweis tested one hypothesis after another: “epidemic influences” (but these would affect both divisions), overcrowding (the Second Division was more crowded), rough examinations, the psychological effect of the priest’s procession with a bell through the First Division (rerouting the priest made no difference), delivery position. None explained the pattern. Then his colleague Jakob Kolletschka was cut by a student’s scalpel during an autopsy and died of an illness with the same symptoms. Semmelweis hypothesized that “cadaveric matter,” carried on the hands of doctors and students who came straight from autopsies, caused the fever. He required handwashing with chlorinated lime, and mortality in the First Division fell dramatically, to around the level of the Second. His explanation was the best: it explained the difference between divisions, Kolletschka’s death, and the effect of the intervention. Tragically, his colleagues largely rejected his findings, partly because they had no theory of germs to explain why it worked, and partly because it implied that doctors were killing their patients. Semmelweis died in an asylum in 1865.
Objections to IBE.
- The “best of a bad lot” problem (Bas van Fraassen, Laws and Symmetry, 1989): IBE can only select the best among the hypotheses we have thought of. If the true explanation isn’t among them, the best available one may be false. Before concluding, ask: have I considered enough alternatives?
- Why is loveliness a guide to truth? Why should the world be simple, unified, or explicable? Some answer that the explanatory virtues have proven reliable in the history of science; others give Bayesian reasons (simpler hypotheses often make sharper predictions, and so gain more when those predictions succeed).
- Stories are seductive. People find explanations satisfying for reasons that have nothing to do with truth: coherence with what they want to believe, narrative appeal, familiarity. The best explanation to us may not be the best explanation of the evidence.
Theory-ladenness of observation
Observations are not simply given. As Norwood Russell Hanson (Patterns of Discovery, 1958) and Thomas Kuhn argued, what we observe depends partly on the concepts and theories we bring to it (see Cognitive penetration and theory-ladenness).
- A trained radiologist sees a tumor in an X-ray where a layperson sees a blur.
- When Galileo showed his colleagues the moons of Jupiter through his telescope, some saw nothing, and some doubted that the telescope gave a true picture of the heavens at all. To accept what the telescope showed, they needed a theory of optics that assured them it did not create illusions.
- Measurements rely on theories of the measuring instruments. A thermometer reading assumes a theory of thermal expansion.
Theory-ladenness is real but has limits. Observations can surprise theories, which could not happen if theory completely determined observation. Anomalies that nobody expected, like the precession of Mercury or the cosmic microwave background, forced changes of theory. And as Jerry Fodor noted, some perceptual processes (as in the Müller-Lyer illusion) are resistant to belief.
Underdetermination
A theory is underdetermined by the evidence if the evidence is also compatible with rival theories.
- Transient underdetermination: the evidence currently available doesn’t decide between rivals, but future evidence may. This is common and unproblematic.
- Permanent underdetermination: two theories are empirically equivalent, so no possible evidence could ever decide between them. Bas van Fraassen gives the example of Newtonian mechanics combined with different hypotheses about the absolute velocity of the center of mass of the universe; all make the same predictions.
The Duhem–Quine thesis implies a kind of underdetermination: since a failed prediction can always be accommodated by revising some auxiliary hypothesis, any theory can in principle be saved.
Does underdetermination mean all theories are equally good? No. Larry Laudan and Jarrett Leplin (“Empirical Equivalence and Underdetermination,” 1991) argued that underdetermination by logic (a theory can be saved) does not imply underdetermination by reasonable standards (a theory should be saved). A theory kept alive only by an ever-growing collection of ad hoc hypotheses is not as good as one that predicts the evidence naturally. Laudan called the view that underdetermination makes all theories equally rational the egalitarian thesis, and argued it is false.
In debates, “the data are consistent with my view too” is a common move. It may be true. But consistency is cheap. The relevant questions are: which view predicted the data rather than accommodating it afterward? Which is simpler? Which requires fewer ad hoc assumptions?
Scientific realism and anti-realism
Do our best scientific theories describe reality, including unobservable entities like electrons, genes, and dark matter? Or are they merely useful instruments for predicting observations?
Scientific realism holds that our best theories are approximately true, and that the unobservable entities they posit really exist.
The main argument for realism is the no-miracles argument (Hilary Putnam, 1975): “Realism is the only philosophy that doesn’t make the success of science a miracle.” If electrons don’t exist, it would be an astonishing coincidence that theories about electrons let us build computers, lasers, and electron microscopes.
The main argument against realism is the pessimistic meta-induction (Larry Laudan, “A Confutation of Convergent Realism,” 1981). The history of science is full of theories that were empirically successful but are now considered false, and whose central terms don’t refer to anything: the crystalline spheres of ancient astronomy, the humoral theory of medicine, phlogiston, caloric (heat as a fluid), the electromagnetic and optical ether, spontaneous generation, vital forces. By induction from this history, our current successful theories will probably turn out false too.
Positions in the debate:
- Constructive empiricism (Bas van Fraassen, The Scientific Image, 1980): the aim of science is empirical adequacy (getting the observable phenomena right), not truth about unobservables. Accepting a theory involves believing only that it is empirically adequate.
- Instrumentalism: theories are tools for prediction and control, not descriptions.
- Structural realism (John Worrall, “Structural Realism: The Best of Both Worlds?,” 1989): what survives theory change is mathematical structure, not claims about the nature of entities. Augustin-Jean Fresnel’s theory of light as vibrations in an elastic ether was abandoned, but his equations were carried over intact into Maxwell’s electromagnetic theory. So we should be realists about the structure our theories describe.
- Entity realism (Ian Hacking, Representing and Intervening, 1983): we are entitled to believe in entities we can manipulate to intervene in other processes, whatever our theories say about them. Of electrons and positrons, which physicists “spray” onto niobium balls to change their charge, Hacking wrote: “if you can spray them, then they are real.”
- Selective realism (Philip Kitcher, Stathis Psillos): be realist about the parts of past theories that were actually responsible for their successes. Those parts have generally been retained.
The debate matters for critical thinking in a subtle way. A healthy attitude toward science combines confidence in well-established findings (the no-miracles intuition) with humility about theoretical details at the frontier (the pessimistic intuition). The observable, predictive core of mature science (vaccines prevent disease, the Earth is billions of years old, CO₂ traps heat) is far more secure than speculative interpretations at the edges.
Causation and causal inference
Many of the most important questions in public life are causal. Does this drug cure the disease? Does this policy reduce crime? Does social media cause depression? Does immigration lower wages? Causal claims are also among the easiest to get wrong.
Hume and the regularity view
David Hume (see Hume) argued that we never observe causal necessity. We observe only that events of one type are constantly conjoined with events of another type, that they are contiguous in space and time, and that the cause comes first. The regularity theory of causation says that causation just is this pattern.
Problems: many regularities are not causal. Night regularly follows day, but day doesn’t cause night. A falling barometer is regularly followed by storms, but the barometer doesn’t cause storms; both are caused by falling air pressure.
Mill’s methods
John Stuart Mill (A System of Logic, 1843) systematized methods for identifying causes.
- Method of agreement: if all cases where the effect occurs share only one factor, that factor is probably the cause. All the wedding guests who got food poisoning ate the prawns.
- Method of difference: if a case where the effect occurs and a case where it doesn’t differ in only one factor, that factor is probably the cause (or part of it). Two guests ate identical meals, except that one ate the prawns and got sick. This is the logic of the controlled experiment.
- Joint method of agreement and difference: combining both.
- Method of residues: subtract the effects of known causes; what remains is due to the remaining factors. The discovery of Neptune used this reasoning: subtract the known planets’ influence on Uranus, and the residual deviation points to an unknown cause.
- Method of concomitant variation: if an effect varies when a factor varies, they are causally connected. The more prawns guests ate, the sicker they were (a dose–response relationship).
Mill’s methods are sound in principle but need care: they assume you have identified all the relevant candidate factors, and real effects often have multiple causes.
Counterfactuals and interventions
The counterfactual theory (David Lewis, “Causation,” 1973): roughly, C caused E if, had C not occurred, E would not have occurred. Aspirin caused my headache to go away if, had I not taken it, the headache would have persisted. This captures the idea behind controlled experiments: compare what happened with what would have happened otherwise.
It faces problems of preemption and overdetermination. Suzy and Billy both throw rocks at a bottle; Suzy’s arrives first and shatters it. Suzy’s throw caused the shattering, but if she hadn’t thrown, Billy’s rock would have shattered it anyway (Ned Hall, 2004). Two assassins shoot a victim at the same instant, each shot fatal: neither shot is necessary, but both seem to be causes.
The interventionist theory (James Woodward, Making Things Happen, 2003; Judea Pearl, Causality, 2000) holds that X causes Y if an intervention that changes X (while leaving everything else that doesn’t depend on X alone) would change Y. This is the idea of causation as a handle for manipulation, and it underlies modern causal inference.
Pearl describes a ladder of causation (The Book of Why, 2018):
- Association (seeing): What does observing X tell me about Y? (“Patients who took the drug recovered more often.”)
- Intervention (doing): What would happen to Y if I made X happen? (“If we give the drug, will patients recover more often?”)
- Counterfactuals (imagining): What would have happened if X had been different? (“Would this patient have recovered without the drug?”)
Statistics alone climbs only the first rung. Climbing higher requires causal assumptions, often represented as diagrams showing which variables affect which.
Correlation and causation
“Correlation does not imply causation” is the most famous slogan in critical thinking. More precisely: when A and B are correlated, there are several possible explanations, and “A causes B” is only one of them.
| Explanation | Structure | Example |
|---|---|---|
| A causes B | A → B | Smoking causes lung cancer. |
| Reverse causation | B → A | Cities with more firefighters at a fire see more damage. The size of the fire causes both; firefighters don’t cause damage. Or: people who are depressed use social media differently, rather than (or as well as) social media causing depression. |
| Confounding (common cause) | C → A and C → B | Ice cream sales and drownings rise together; hot weather causes both. |
| Chance | none | Across years, the number of films Nicolas Cage appeared in correlated with the number of people who drowned by falling into swimming pools (from Tyler Vigen’s Spurious Correlations). Search enough variables and some will correlate by coincidence. |
| Selection effects / collider bias | A → S ← B, and we only look at S | Among hospitalized patients, two unrelated diseases can appear negatively correlated, because having either one can get you admitted (Berkson’s paradox, 1946). |
| Bidirectional causation | A ⇄ B | Poverty and poor health each worsen the other. |
The HRT lesson. For decades, large observational studies such as the Nurses’ Health Study found that women taking hormone replacement therapy (HRT) after menopause had lower rates of heart disease. Many doctors recommended HRT partly for heart protection. Then the Women’s Health Initiative, a large randomized controlled trial, reported in 2002 that combined HRT did not protect against heart disease, and in fact slightly increased the risk of heart attacks, strokes, and breast cancer. The observational association had been largely due to confounding: women who chose to take HRT were, on average, healthier, wealthier, and more health-conscious (the healthy user effect). (The picture has since become more nuanced, with the timing of HRT relative to menopause appearing to matter, but the core lesson stands.)
Sunscreen and skin cancer. Some studies have found that people who use more sunscreen have higher rates of melanoma. Does sunscreen cause cancer? No: people who spend more time in intense sun use more sunscreen. Sun exposure confounds the association.
Checklist for a causal claim based on correlation:
- Could B cause A instead?
- What third factor could cause both? Were such factors measured and controlled for?
- Could the correlation be due to chance? How many variables were examined?
- How were the cases selected? Could the selection itself create the correlation?
- Is there a dose–response relationship? Did the cause come first?
- Is there a plausible mechanism?
- Has the finding been confirmed by experiments or natural experiments?
Randomized controlled trials and the hierarchy of evidence
The best tool for establishing causation is the randomized controlled trial (RCT).
- Control group: compare people who receive the treatment with similar people who don’t.
- Randomization: assign people to groups by chance. This breaks the link between the treatment and every possible confounder, known or unknown. On average, the groups will be alike in age, health, wealth, motivation, genetics, and everything else, so any difference in outcomes can be attributed to the treatment.
- Blinding: participants (and ideally researchers and assessors) don’t know who received the treatment, preventing placebo effects and observer bias from creating differences.
Historical milestones: James Lind’s 1747 comparison of six treatments for scurvy on board HMS Salisbury, which found that citrus fruit worked (controlled, but not randomized); and the British Medical Research Council’s 1948 trial of streptomycin for tuberculosis, designed with the statistician Austin Bradford Hill, often cited as the first modern randomized controlled trial.
The hierarchy of evidence, as commonly taught in evidence-based medicine (from strongest to weakest, for questions about treatment effects):
- Systematic reviews and meta-analyses of randomized controlled trials.
- Randomized controlled trials.
- Cohort studies (following groups with and without an exposure over time).
- Case-control studies (comparing people with a condition to people without it, looking back for differences in exposure).
- Case series and case reports.
- Expert opinion, mechanistic reasoning, and anecdote.
The hierarchy is a useful default, not an absolute rule:
- Some questions can’t be answered by RCTs. You can’t randomize people to smoke for 30 years. A famous satirical paper (Gordon Smith and Jill Pell, “Parachute Use to Prevent Death and Major Trauma Related to Gravitational Challenge: Systematic Review of Randomised Controlled Trials,” BMJ, 2003) found no RCTs of parachutes and concluded, tongue in cheek, that their effectiveness was unproven. When an effect is enormous and obvious, observational evidence can suffice.
- External validity. In 2018, researchers actually ran a randomized trial of parachutes (Robert Yeh and colleagues, BMJ): participants jumped from aircraft with or without parachutes, and there was no difference in death or injury. The aircraft were parked on the ground, and participants jumped about 0.6 meters. The trial was internally valid, and useless for the question that matters. The satire makes a serious point: a trial shows what worked for its population, in its conditions. Nancy Cartwright and Jeremy Hardie (Evidence-Based Policy, 2012) stress that “it worked there” does not guarantee “it will work here” unless the causal factors that made it work are also present here.
- Natural experiments can approach the strength of RCTs. In London in 1854, John Snow traced a cholera outbreak to the Broad Street water pump. In his “Grand Experiment,” he compared households supplied by the Southwark and Vauxhall water company, which drew water from sewage-contaminated parts of the Thames, with households supplied by the Lambeth company, which had moved its intake upstream. The households were intermixed in the same neighborhoods, similar in other respects, and mostly unaware of which company supplied them. Deaths from cholera were many times higher among Southwark and Vauxhall customers.
The Bradford Hill criteria
How do you establish causation when RCTs are impossible? In 1965, Austin Bradford Hill (“The Environment and Disease: Association or Causation?”) listed nine “viewpoints” for judging whether an observed association is causal:
- Strength: large effects are less likely to be due to confounding. (Heavy smokers had many times the lung cancer risk of non-smokers.)
- Consistency: the association appears in different populations, places, and study designs.
- Specificity: the exposure is associated with a specific disease (weakest criterion).
- Temporality: the cause precedes the effect. (The only essential criterion.)
- Biological gradient: dose–response. More exposure, more effect.
- Plausibility: a plausible mechanism exists.
- Coherence: the causal interpretation fits with what is known about the disease.
- Experiment: removing the exposure reduces the effect (e.g., people who quit smoking have lower risk).
- Analogy: similar causes have similar effects.
Hill insisted these were not a checklist: “None of my nine viewpoints can bring indisputable evidence for or against the cause-and-effect hypothesis and none can be required as a sine qua non.”
Smoking and lung cancer is the great example. Richard Doll and Bradford Hill’s case-control study (1950) and their British Doctors Study (begun in 1951) established the link with strength, consistency, temporality, and dose–response, without any randomized trial. The tobacco industry’s defense was “correlation is not causation.” Remarkably, the great statistician R. A. Fisher (himself a smoker, and a paid consultant to the tobacco industry) argued in the late 1950s that a genetic factor might cause both a craving for smoking and a susceptibility to cancer. The hypothesis was not absurd in principle, but it could not account for the full pattern of evidence, such as the falling risk among people who quit. Even brilliant experts can be wrong, especially where their interests or habits are involved.
Values in science
Is science value-free? Scientists should not let their political or personal preferences decide what the evidence shows. But philosophers have argued that values enter science in legitimate ways.
Inductive risk. Richard Rudner (“The Scientist Qua Scientist Makes Value Judgments,” 1953) argued that no scientific hypothesis is ever completely verified. So in accepting a hypothesis, a scientist must decide that the evidence is sufficiently strong, and “how sure we need to be before we accept a hypothesis will depend on how serious a mistake would be.” We would demand much stronger evidence before concluding that a drug ingredient is not toxic than before concluding that a batch of belt buckles is not defective. Carl Hempel (1965) called this inductive risk. Heather Douglas (Science, Policy, and the Value-Free Ideal, 2009) developed the argument: values legitimately play an indirect role in deciding how much evidence is enough, especially where the consequences of error are serious, but they should not play a direct role, treating a preferred conclusion as a reason to believe it.
Objectivity as a social achievement. Helen Longino (Science as Social Knowledge, 1990) argued that individual scientists cannot be free of bias, but a scientific community can achieve objectivity if it has: (1) recognized venues for criticism, (2) uptake of criticism (the community actually changes in response), (3) shared public standards, and (4) “tempered equality of intellectual authority” among qualified members, including those with different perspectives.
Epistemic values, such as Kuhn’s accuracy, consistency, scope, simplicity, and fruitfulness, are values too. They help decide between theories, and they are part of what makes science rational.
Conflicts of interest. Funding sources can affect results. A Cochrane systematic review (Andreas Lundh and colleagues, 2017) found that drug and device studies sponsored by manufacturers more often had favorable efficacy results and conclusions than studies with other sponsors. This does not mean industry-funded research is worthless. It means that the source of funding is relevant background information, and that independent replication is especially important.
Scientific consensus
A scientific consensus exists when the great majority of experts in a field agree on a conclusion, usually as a result of many studies converging. For example, surveys of the climate science literature have found that the great majority of papers expressing a position on the cause of recent global warming attribute it to human activity (John Cook and colleagues, 2013, reported 97% of such abstracts). National academies of science around the world have issued statements to the same effect.
Why consensus matters to non-experts. Most of us cannot evaluate the primary evidence in climatology, immunology, or particle physics. The existence of a consensus is itself strong evidence, provided it arises for the right reasons. Boaz Miller (“When Is Consensus Knowledge-Based?,” 2013) argues that consensus is a good indicator of knowledge when it rests on apparent consilience of evidence (multiple independent lines of evidence point the same way), social calibration (experts share standards of evidence and interpretation), and social diversity (the consensus includes people with different backgrounds, interests, and methods, so it isn’t the product of a shared bias).
Consensus has been wrong. Alfred Wegener’s continental drift was rejected for about half a century. The Australian doctors Barry Marshall and Robin Warren proposed in the early 1980s that most stomach ulcers were caused by the bacterium Helicobacter pylori, not by stress or acid. They met widespread skepticism. Marshall famously drank a culture of the bacterium in 1984, developed gastritis, and treated it with antibiotics. They received the Nobel Prize in 2005.
What follows? Two things:
- The base rate matters. Mature scientific consensus is overturned much less often than it is confirmed, and when it is overturned, the old view usually survives as an approximation (Newtonian mechanics is still used to send spacecraft to other planets).
- The corrections came from within science, through evidence. Wegener was vindicated by oceanographic data, Marshall by experiments and trials. The lesson of these episodes is “follow the evidence,” not “ignore the consensus.”
The Galileo gambit (“They laughed at Galileo, and he was right, so they’re laughing at me and I must be right”) ignores this. As Carl Sagan put it in Broca’s Brain (1979): “The fact that some geniuses were laughed at does not imply that all who are laughed at are geniuses. They laughed at Columbus, they laughed at Fulton, they laughed at the Wright brothers. But they also laughed at Bozo the Clown.” See Experts and novices.
The replication crisis
Since the early 2010s, many fields have discovered that a large share of published findings do not replicate when independent researchers repeat the studies.
- The Open Science Collaboration (“Estimating the Reproducibility of Psychological Science,” Science, 2015) attempted to replicate 100 studies published in major psychology journals. Of the original studies, 97% had reported significant results; only about 36% of the replications did, and the average effect size in the replications was about half the original.
- In cancer biology, scientists at the biotechnology company Amgen reported (C. Glenn Begley and Lee Ellis, Nature, 2012) that they could confirm the findings of only 6 of 53 “landmark” preclinical studies. Researchers at Bayer reported similar problems (Florian Prinz and colleagues, 2011).
- Some famous findings in psychology have failed large-scale replications, including ego depletion (the idea that willpower is a limited resource that gets “used up”; a large multi-lab replication in 2016 found an effect near zero) and power posing (the claim that adopting expansive postures raises testosterone and risk-taking; one of the original authors, Dana Carney, publicly stated in 2016 that she no longer believed the effect was real).
Causes (see also Statistics and significance):
- Publication bias (the file drawer problem, as Robert Rosenthal called it in 1979): studies with significant results are more likely to be published; null results end up in the file drawer. The published literature therefore over-represents positive findings.
- P-hacking and the garden of forking paths.
- Low statistical power: small studies that can only detect effects if they are exaggerated by chance.
- HARKing (Hypothesizing After the Results are Known; Norbert Kerr, 1998): presenting a post hoc finding as if it had been predicted.
- Incentives: careers depend on publishing novel, positive results.
- Fraud, which is rare but real: the social psychologist Diederik Stapel fabricated data in dozens of papers, discovered in 2011; Andrew Wakefield’s 1998 paper claiming a link between the MMR vaccine and autism was retracted in 2010 after investigations found serious misconduct.
Reforms include pre-registration of studies, registered reports (journals accept papers for publication based on their methods, before results are known), open data and code, larger samples, and large multi-lab replication projects.
Pseudoscience and how to spot it
No single criterion separates science from pseudoscience, but the warning signs cluster. Adapted from Carl Sagan, Mario Bunge, and Scott Lilienfeld and colleagues:
- Unfalsifiability. The claim is compatible with every possible outcome.
- Ad hoc rescues. Failures are explained away with untestable excuses (“the energy was disturbed,” “you didn’t believe enough”).
- Reliance on testimonials and anecdotes rather than controlled studies.
- No self-correction. The doctrine is essentially unchanged after decades or centuries, despite new evidence. Astrology is still based on a system set out in antiquity.
- Reversal of the burden of proof. “You can’t prove it doesn’t work.” See Burden of proof.
- Scientific-sounding jargon without precise meaning (“quantum healing,” “energy fields,” “toxins,” “vibrations”).
- Isolation from other sciences. The claims don’t connect with, and often contradict, well-established physics, chemistry, and biology.
- Extraordinary claims on weak evidence.
- The persecution narrative. “The establishment is suppressing this” (see the Galileo gambit).
- Emphasis on confirmation rather than attempted refutation.
- Avoidance of peer review and independent testing.
- Financial incentives tied to belief (products, courses, consultations).
- Appeals to ancient wisdom or to nature as guarantees of truth.
Homeopathy illustrates several of these. It holds that substances causing symptoms in healthy people cure similar symptoms in the sick, and that dilution increases potency. A common homeopathic dilution, “30C,” means diluting by a factor of 100, thirty times over: a total dilution of 10⁶⁰. For comparison, a mole of a substance contains about 6 × 10²³ molecules. At 30C, the chance that a dose contains even a single molecule of the original substance is effectively zero. The claimed mechanism contradicts basic chemistry, and systematic reviews have not found effects beyond placebo (for example, Australia’s National Health and Medical Research Council concluded in 2015 that there were no health conditions for which there was reliable evidence that homeopathy was effective).
Astrology has been tested. In a double-blind study published in Nature (Shawn Carlson, “A Double-Blind Test of Astrology,” 1985), astrologers were unable to match natal charts to the correct personality profiles better than chance.
Part of astrology’s apparent success is the Forer effect (or Barnum effect). In 1948 the psychologist Bertram Forer gave his students a personality test and then gave each of them a “personalized” description. In fact, every student received the same description, assembled largely from a newsstand astrology book (“You have a great need for other people to like and admire you... At times you are extroverted, affable, sociable, while at other times you are introverted, wary, reserved...”). Students rated its accuracy, on average, at about 4.3 on a scale of 0 to 5. Vague, flattering, and double-sided statements feel personally accurate to almost everyone.
Check your understanding
1Explain the asymmetry between verification and falsification, and why it isn’t the whole story.
Answer
No number of confirming instances proves a universal law, but one counterexample logically refutes it. It isn’t the whole story because of the Duhem–Quine problem: predictions depend on auxiliary assumptions, so a failed prediction may refute an auxiliary rather than the theory. Also, probabilistic and existential hypotheses aren’t strictly falsifiable, and scientists reasonably retain theories in the face of some anomalies.
2When Uranus’s orbit didn’t match Newtonian predictions, astronomers postulated a new planet. Was this ad hoc? Why or why not?
Answer
No. The hypothesis was independently testable: it predicted a planet of a certain mass at a certain position, which could be looked for, and it was found (Neptune, 1846). An ad hoc hypothesis has no consequences beyond saving the theory.
3What distinguishes a progressive from a degenerating research programme?
Answer
A progressive programme’s modifications predict novel facts, some of which are confirmed. A degenerating programme’s modifications only accommodate known facts after the event, and its predictions fail or lag behind the evidence.
4“Children who eat breakfast do better at school, so schools should provide breakfast to improve grades.” Give three alternative explanations for the correlation.
Answer
(1) Confounding: families with more resources and stability both provide breakfast and support schoolwork. (2) Confounding by health or sleep: well-rested, healthy children are more likely to eat breakfast and to concentrate. (3) Reverse causation or selection: motivated, organized children may be more likely to get up in time for breakfast. The claim may still be true (and there is some experimental evidence on school breakfast programs), but the correlation alone doesn’t establish it.
5Why does randomization matter in a controlled trial?
Answer
It makes the treatment and control groups similar, on average, in all respects, including factors nobody thought to measure. This breaks any systematic link between the treatment and confounders, so differences in outcome can be attributed to the treatment (up to chance, which statistics can quantify).
6What are the no-miracles argument and the pessimistic meta-induction?
Answer
No-miracles: the success of science would be a miracle unless our theories were approximately true, so they probably are. Pessimistic meta-induction: many past successful theories were false (phlogiston, ether, caloric), so our current successful theories are probably false too. Structural realism and selective realism try to accommodate both.
7Someone says: “Scientists were wrong about ulcers and continental drift, so I don’t trust the scientific consensus on vaccines.” Evaluate.
Answer
The premise is true, but the conclusion doesn’t follow. Consensus is sometimes overturned, but much less often than it is confirmed; the base rate matters. The corrections came from new evidence within science, not from rejecting science. The consensus on vaccine safety rests on many independent lines of evidence (large observational studies of millions, trials, surveillance systems in many countries). The appropriate response to fallibility is to follow the evidence, which here strongly supports the consensus, not to reject it wholesale. There is also an asymmetry test: would the speaker apply the same reasoning to reject consensus claims they like?
8Why is the Forer effect relevant to evaluating personality tests, astrology, and “cold reading” by psychics?
Answer
Because vague, general, and double-sided descriptions feel personally accurate to almost everyone. A subjective sense of “that’s so me” is therefore no evidence that the method has any real diagnostic power. The proper test is whether people can pick out their own description from others’ better than chance, under blind conditions.
Further reading
Introductions
- Samir Okasha, Philosophy of Science: A Very Short Introduction (Oxford University Press, 2nd ed. 2016). The best short introduction.
- A. F. Chalmers, What Is This Thing Called Science? (Hackett / Open University Press, 4th ed. 2013).
- Peter Godfrey-Smith, Theory and Reality: An Introduction to the Philosophy of Science (University of Chicago Press, 2nd ed. 2021).
Classics
- Karl Popper, Conjectures and Refutations (Routledge, 1963), ch. 1, “Science: Conjectures and Refutations.”
- Thomas Kuhn, The Structure of Scientific Revolutions (University of Chicago Press, 1962; 50th anniversary ed. 2012 with an introduction by Ian Hacking).
- Imre Lakatos, “Falsification and the Methodology of Scientific Research Programmes” (1970), in Lakatos and Musgrave (eds.), Criticism and the Growth of Knowledge.
- Carl Hempel, Philosophy of Natural Science (Prentice-Hall, 1966). Short and clear; source of the Semmelweis example.
- Peter Lipton, Inference to the Best Explanation (Routledge, 2nd ed. 2004).
Science and society
- Naomi Oreskes, Why Trust Science? (Princeton University Press, 2019).
- Heather Douglas, Science, Policy, and the Value-Free Ideal (University of Pittsburgh Press, 2009).
- Stuart Ritchie, Science Fictions: Exposing Fraud, Bias, Negligence and Hype in Science (Bodley Head, 2020). On the replication crisis.
- Ben Goldacre, Bad Science (Fourth Estate, 2008) and Bad Pharma (2012).
- Carl Sagan, The Demon-Haunted World: Science as a Candle in the Dark (Random House, 1995).
Causal inference
- Judea Pearl and Dana Mackenzie, The Book of Why (Basic Books, 2018).
- Nancy Cartwright and Jeremy Hardie, Evidence-Based Policy: A Practical Guide to Doing It Better (Oxford University Press, 2012).
- Austin Bradford Hill, “The Environment and Disease: Association or Causation?,” Proceedings of the Royal Society of Medicine 58 (1965). Short and readable.
Concepts from this chapter
Each has its own page with the key idea, objections and replies, common mistakes, and a self-check, in English and Persian.
