Mastering EpistemologyGuide · Map · Audio فا
An Experiment on a Bird in the Air Pump by Joseph Wright of Derby (1768)

The garden of forking paths

باغِ راه‌های چنگالی

Andrew Gelman & Eric Loken
Joseph Wright of Derby, An Experiment on a Bird in the Air Pump, 1768

The garden of forking paths shows how many data-dependent but individually defensible analytic choices can inflate false positives even without conscious p-hacking.

1The problem it solves

Classical p-values assume the analysis rule was fixed independently of the realized data. In practice, researchers decide exclusions, transformations, outcomes, interactions, covariates, and stopping rules after seeing partial patterns. The chosen path may look innocent, but the counterfactual set of paths means the nominal error rate no longer describes the procedure actually used.

2The idea

Even with no intent to cheat, analysts make defensible choices that depend on how the data happened to look. The test was therefore never truly pre-specified, and the false positive rate is far above the nominal one.

3Origin & context

Andrew Gelman and Eric Loken popularized the “garden of forking paths” to distinguish researcher degrees of freedom from deliberate multiple-testing misconduct. It is a structural argument for pre-specification, multiverse analysis, and transparency.

4Canonical example

You would have excluded different outliers if the data had looked different. That counterfactual is enough to break the test.

5Objections & replies

Objection. If each choice is scientifically reasonable, there is no reason to penalize the analysis.

Reply. Reasonableness of each branch does not restore the advertised sampling distribution when the branch selected depends on the data. The inferential procedure includes the selection rule, not just the final regression.

Objection. Pre-registration is too rigid for exploratory science.

Reply. Exploration is valuable; the solution is labeling and separation. Explore freely, then treat resulting hypotheses as generated rather than independently confirmed, or validate them on new data.

6Don't confuse it with

Forking paths vs. p-hacking

P-hacking usually implies intentional or at least goal-directed search for significance. Forking paths can occur in complete good faith through data-responsive decisions.

Multiple comparisons vs. forking paths

Multiple-comparison corrections handle explicit families of tests; forking paths include implicit tests that were never run because choices depended on earlier observations.

7Common mistakes

  • Defending a result with “we only ran one final model” when many alternative models were implicitly available.
  • Treating preregistration as a ban on exploration rather than a way to distinguish exploration from confirmation.

8In your work

The argument for pre-registration even when you have no intention of cheating. Good faith is not a defence, because the mechanism does not require bad faith.

9Check yourself

You would have used a log transform if the raw outcome had looked more skewed, but it did not, so you report the raw analysis. Why can that still affect inference?

Show answer

Because the analysis rule depended on the observed data. The counterfactual branch is part of the procedure, so the nominal p-value for a fixed raw analysis does not fully describe the selection process.

«باغِ مسیرهای منشعب» نشان می‌دهد مجموعه‌ای از انتخاب‌های وابسته به داده که هر کدام جداگانه قابل‌دفاع‌اند می‌توانند بدون هیچ p-hacking عمدی نرخِ مثبتِ کاذب را بالا ببرند.

1مسئله‌ای که حل می‌کند

مقدارِ p کلاسیک فرض می‌کند قاعدهٔ تحلیل مستقل از دادهٔ مشاهده‌شده از قبل ثابت بوده است. در عمل پژوهشگر حذف، تبدیل، پیامد، تعامل، کنترل و توقف را با دیدنِ الگوها انتخاب می‌کند. مسیرِ نهایی شاید معقول باشد، اما مجموعهٔ مسیرهای خلاف‌واقعی یعنی نرخِ اسمی دیگر روشِ واقعی را توصیف نمی‌کند.

2ایدهٔ اصلی

حتی بدون قصدِ تقلب، تحلیل‌گران انتخاب‌هایی قابل دفاع می‌کنند که به ظاهرِ اتفاقیِ داده وابسته است. پس آزمون هرگز واقعاً از پیش تعیین نشده بود و نرخِ مثبتِ کاذب بسیار بالاتر از نرخِ اسمی است.

3خاستگاه و زمینه

اندرو گلمان و اریک لوکن اصطلاحِ «باغِ مسیرهای منشعب» را برای جداکردنِ آزادیِ تحلیلی از دستکاریِ عمدیِ آزمون‌ها مشهور کردند. استدلالی ساختاری به سودِ پیش‌ثبت، تحلیلِ چندجهانی و شفافیت است.

4مثالِ کلاسیک

اگر داده جور دیگری بود، پرت‌های دیگری را کنار می‌گذاشتید. همین خلافِ واقع برای شکستنِ آزمون کافی است.

5اعتراض‌ها و پاسخ‌ها

اعتراض. اگر هر انتخاب علمی معقول است، چرا باید تحلیل جریمه شود؟

پاسخ. معقول‌بودنِ هر شاخه توزیعِ نمونه‌گیریِ اعلام‌شده را برنمی‌گرداند وقتی انتخابِ شاخه به داده وابسته است. روشِ استنتاج شامل قاعدهٔ انتخاب هم هست، نه فقط رگرسیونِ نهایی.

اعتراض. پیش‌ثبت برای علمِ اکتشافی بیش‌ازحد سخت است.

پاسخ. اکتشاف ارزشمند است؛ راه‌حل برچسب‌گذاری و جداسازی است. آزادانه کشف کنید، اما فرضیهٔ حاصل را تولیدشده بدانید نه تأییدشده، یا روی دادهٔ تازه اعتبارسنجی کنید.

6با این‌ها اشتباه نگیرید

مسیرهای منشعب و p-hacking

p-hacking معمولاً جست‌وجوی هدفمند برای معناداری است؛ مسیرهای منشعب می‌تواند کاملاً با حسن‌نیت از تصمیم‌های پاسخ‌گو به داده ایجاد شود.

چندآزمونی و مسیرهای منشعب

اصلاحِ چندآزمونی خانوادهٔ آزمون‌های صریح را می‌گیرد؛ مسیرهای منشعب آزمون‌های ضمنی‌ای را هم شامل می‌شود که اصلاً اجرا نشدند چون انتخاب‌ها به مشاهدهٔ قبلی وابسته بودند.

7خطاهای رایج

  • دفاع با «فقط یک مدلِ نهایی اجرا کردیم» وقتی مدل‌های جایگزین به‌صورتِ ضمنی در دسترس بودند.
  • فهمِ پیش‌ثبت به‌عنوان ممنوعیتِ اکتشاف به‌جای راهی برای جداکردنِ اکتشاف و تأیید.

8در کارِ شما

استدلالی برای ثبتِ پیشاپیشِ طرح، حتی وقتی قصدِ تقلب ندارید. حسنِ نیت دفاع نیست، چون این سازوکار نیازی به سوءنیت ندارد.

9خودآزمایی

اگر پیامد کج‌تر بود آن را لگاریتمی می‌کردید، اما چون نبود تحلیلِ خام را گزارش کردید. چرا همین تصمیمِ اجرا‌نشده هم ممکن است استنتاج را متاثر کند؟

نمایش پاسخ

چون قاعدهٔ تحلیل به دادهٔ مشاهده‌شده وابسته بود. شاخهٔ خلاف‌واقعی بخشی از روش است و مقدارِ p برای تحلیلِ خامِ ازپیش‌ثابت کلِ فرایندِ انتخاب را توصیف نمی‌کند.