Mastering EpistemologyGuide · Map · Audio فا
An Experiment on a Bird in the Air Pump by Joseph Wright of Derby (1768)

The replication crisis

بحرانِ تکرارپذیری

Ioannidis · Meehl · Gelman · Nosek
Joseph Wright of Derby, An Experiment on a Bird in the Air Pump, 1768

The replication crisis is evidence that published literatures can systematically overstate reliability because incentives and analytic flexibility filter which findings become visible and stable.

1The problem it solves

Failure to reproduce famous results exposed a structural rather than merely individual problem. Low power, publication bias, flexible analysis, weak measurement, small samples, and novelty incentives can generate a literature where many “successful” papers are expected to fail independent repetition. Replication changes the prior on isolated claims and motivates institutional reform.

2The idea

A large share of published findings fail when independently repeated. The causes are mostly structural — incentives, flexible analysis, selective publication — rather than fraudulent.

3Origin & context

The crisis became prominent through large replication projects in psychology and medicine, meta-research associated with figures such as John Ioannidis, Brian Nosek, Andrew Gelman, and earlier methodological warnings from Paul Meehl and others. It is not one event but a cluster of reliability problems.

4Canonical example

Whole subfields of psychology and medicine failed systematic replication attempts, including results taught as settled.

5Objections & replies

Objection. Failure to reproduce exactly does not mean the original result was false; context and heterogeneity matter.

Reply. Correct. Replication must distinguish direct reproduction from conceptual generalization, and failed replication can arise from genuine moderation. But unexplained fragility still weakens broad claims of robustness.

Objection. Science is self-correcting, so the crisis demonstrates success rather than failure.

Reply. Both can be true. Detection and reform are self-correction; the scale of the problem shows that correction can be slow and that publication systems systematically generated overconfidence before correction.

6Don't confuse it with

Replication vs. reproducibility

Reproducibility often means obtaining the same result from the same data/code; replication usually means testing the claim with new data or a new study.

Failed replication vs. falsification

A replication failure can implicate sampling, measurement, context, or analysis as well as the substantive hypothesis. It is evidence, not automatically a logically decisive refutation.

7Common mistakes

  • Treating every published result as equally suspect instead of using design quality, preregistration, sample size, and independent replication to update differentially.
  • Treating one failed replication as proof of fraud.

8In your work

Sets a prior on any single published result — including, and especially, the ones that support your position.

9Check yourself

A surprising single-study finding has p=0.02 but no preregistration, low power, and no independent replication. How should the replication crisis affect your prior?

Show answer

It should materially lower confidence relative to the headline result. The relevant base rate includes selective publication and analytic flexibility; independent, well-powered replication should carry much more weight.

بحرانِ تکرار نشان می‌دهد ادبیاتِ چاپ‌شده می‌تواند نظام‌مند قابلیتِ اتکا را بیش‌برآورد کند، چون انگیزه و انعطافِ تحلیل تعیین می‌کنند کدام یافته دیده و پایدار می‌شود.

1مسئله‌ای که حل می‌کند

ناتوانی در بازتولیدِ نتایجِ مشهور مسئله‌ای ساختاری را آشکار کرد، نه فقط خطای فردی. توانِ کم، سوگیریِ انتشار، آزادیِ تحلیل، سنجشِ ضعیف، نمونهٔ کوچک و انگیزهٔ تازگی می‌توانند ادبیاتی بسازند که بسیاری از «موفقیت‌ها» در تکرارِ مستقل شکست بخورند. این بحران پیشینِ ما دربارهٔ یافتهٔ منفرد را عوض می‌کند.

2ایدهٔ اصلی

بخش بزرگی از یافته‌های منتشرشده هنگام تکرارِ مستقل شکست می‌خورند. علت‌ها بیشتر ساختاری‌اند — انگیزه‌ها، تحلیلِ انعطاف‌پذیر، انتشارِ گزینشی — نه تقلب.

3خاستگاه و زمینه

با پروژه‌های بزرگِ تکرار در روان‌شناسی و پزشکی و فراتحقیقِ چهره‌هایی چون جان یوانیدیس، برایان نوزک، اندرو گلمان و هشدارهای قدیمی‌ترِ پل میل برجسته شد. یک رویدادِ واحد نیست، بلکه خوشه‌ای از مشکلاتِ وثاقت است.

4مثالِ کلاسیک

زیررشته‌های کاملی از روان‌شناسی و پزشکی در تلاش‌های نظام‌مندِ تکرار شکست خوردند، از جمله نتایجی که به‌عنوان امرِ قطعی تدریس می‌شد.

5اعتراض‌ها و پاسخ‌ها

اعتراض. شکستِ بازتولید دقیقاً یعنی نتیجهٔ اصلی کاذب نبود؛ زمینه و ناهمگنی مهم‌اند.

پاسخ. درست. باید تکرارِ مستقیم را از تعمیمِ مفهومی جدا کرد و تعدیلِ واقعی ممکن است وجود داشته باشد. اما شکنندگیِ توضیح‌نداده‌شده ادعای کلیِ استحکام را ضعیف می‌کند.

اعتراض. علم خوداصلاح است، پس بحران موفقیتِ علم را نشان می‌دهد نه شکست را.

پاسخ. هر دو می‌توانند درست باشند. کشف و اصلاح خوداصلاحی است؛ مقیاسِ مشکل نشان می‌دهد اصلاح می‌تواند کند باشد و نظامِ انتشار پیش از اصلاح بیش‌اطمینانی تولید کند.

6با این‌ها اشتباه نگیرید

تکرار و بازتولیدپذیری

بازتولیدپذیری اغلب یعنی گرفتنِ همان نتیجه از همان داده/کد؛ تکرار یعنی آزمونِ ادعا با داده یا مطالعهٔ تازه.

شکستِ تکرار و ابطال

شکست می‌تواند از نمونه، سنجش، زمینه یا تحلیل باشد و فقط فرضیهٔ substantive را هدف نگیرد. شاهد است، نه الزاماً ابطالِ منطقیِ یک‌ضرب.

7خطاهای رایج

  • یکسان مشکوک دانستنِ همهٔ مقالات به‌جای وزن‌دهی بر اساسِ طراحی، پیش‌ثبت، اندازهٔ نمونه و تکرارِ مستقل.
  • برداشتِ یک شکستِ تکرار به‌عنوان اثباتِ تقلب.

8در کارِ شما

احتمالِ پیشینی برای هر نتیجهٔ منتشرشدهٔ منفرد تعیین می‌کند — از جمله و به‌ویژه آن‌هایی که موضعِ شما را تأیید می‌کنند.

9خودآزمایی

یافته‌ای شگفت‌آور p=0.02 دارد اما پیش‌ثبت نشده، توانِ کم و هیچ تکرارِ مستقلی ندارد. بحرانِ تکرار چه اثری بر پیشین شما دارد؟

نمایش پاسخ

باید اعتماد را به‌طور معنادار پایین‌تر از تیتر نگه دارد. نرخِ پایه شامل انتخابِ انتشار و آزادیِ تحلیل است؛ تکرارِ مستقل و پرتوان باید وزنِ بسیار بیشتری بگیرد.