The garden of forking paths shows how many data-dependent but individually defensible analytic choices can inflate false positives even without conscious p-hacking.
1The problem it solves
Classical p-values assume the analysis rule was fixed independently of the realized data. In practice, researchers decide exclusions, transformations, outcomes, interactions, covariates, and stopping rules after seeing partial patterns. The chosen path may look innocent, but the counterfactual set of paths means the nominal error rate no longer describes the procedure actually used.
2The idea
Even with no intent to cheat, analysts make defensible choices that depend on how the data happened to look. The test was therefore never truly pre-specified, and the false positive rate is far above the nominal one.
3Origin & context
Andrew Gelman and Eric Loken popularized the “garden of forking paths” to distinguish researcher degrees of freedom from deliberate multiple-testing misconduct. It is a structural argument for pre-specification, multiverse analysis, and transparency.
4Canonical example
You would have excluded different outliers if the data had looked different. That counterfactual is enough to break the test.
5Objections & replies
Objection. If each choice is scientifically reasonable, there is no reason to penalize the analysis.
Reply. Reasonableness of each branch does not restore the advertised sampling distribution when the branch selected depends on the data. The inferential procedure includes the selection rule, not just the final regression.
Objection. Pre-registration is too rigid for exploratory science.
Reply. Exploration is valuable; the solution is labeling and separation. Explore freely, then treat resulting hypotheses as generated rather than independently confirmed, or validate them on new data.
6Don't confuse it with
Forking paths vs. p-hacking
P-hacking usually implies intentional or at least goal-directed search for significance. Forking paths can occur in complete good faith through data-responsive decisions.
Multiple comparisons vs. forking paths
Multiple-comparison corrections handle explicit families of tests; forking paths include implicit tests that were never run because choices depended on earlier observations.
7Common mistakes
- Defending a result with “we only ran one final model” when many alternative models were implicitly available.
- Treating preregistration as a ban on exploration rather than a way to distinguish exploration from confirmation.
8In your work
The argument for pre-registration even when you have no intention of cheating. Good faith is not a defence, because the mechanism does not require bad faith.
9Check yourself
You would have used a log transform if the raw outcome had looked more skewed, but it did not, so you report the raw analysis. Why can that still affect inference?
Show answer
Because the analysis rule depended on the observed data. The counterfactual branch is part of the procedure, so the nominal p-value for a fixed raw analysis does not fully describe the selection process.
«باغِ مسیرهای منشعب» نشان میدهد مجموعهای از انتخابهای وابسته به داده که هر کدام جداگانه قابلدفاعاند میتوانند بدون هیچ p-hacking عمدی نرخِ مثبتِ کاذب را بالا ببرند.
1مسئلهای که حل میکند
مقدارِ p کلاسیک فرض میکند قاعدهٔ تحلیل مستقل از دادهٔ مشاهدهشده از قبل ثابت بوده است. در عمل پژوهشگر حذف، تبدیل، پیامد، تعامل، کنترل و توقف را با دیدنِ الگوها انتخاب میکند. مسیرِ نهایی شاید معقول باشد، اما مجموعهٔ مسیرهای خلافواقعی یعنی نرخِ اسمی دیگر روشِ واقعی را توصیف نمیکند.
2ایدهٔ اصلی
حتی بدون قصدِ تقلب، تحلیلگران انتخابهایی قابل دفاع میکنند که به ظاهرِ اتفاقیِ داده وابسته است. پس آزمون هرگز واقعاً از پیش تعیین نشده بود و نرخِ مثبتِ کاذب بسیار بالاتر از نرخِ اسمی است.
3خاستگاه و زمینه
اندرو گلمان و اریک لوکن اصطلاحِ «باغِ مسیرهای منشعب» را برای جداکردنِ آزادیِ تحلیلی از دستکاریِ عمدیِ آزمونها مشهور کردند. استدلالی ساختاری به سودِ پیشثبت، تحلیلِ چندجهانی و شفافیت است.
4مثالِ کلاسیک
اگر داده جور دیگری بود، پرتهای دیگری را کنار میگذاشتید. همین خلافِ واقع برای شکستنِ آزمون کافی است.
5اعتراضها و پاسخها
اعتراض. اگر هر انتخاب علمی معقول است، چرا باید تحلیل جریمه شود؟
پاسخ. معقولبودنِ هر شاخه توزیعِ نمونهگیریِ اعلامشده را برنمیگرداند وقتی انتخابِ شاخه به داده وابسته است. روشِ استنتاج شامل قاعدهٔ انتخاب هم هست، نه فقط رگرسیونِ نهایی.
اعتراض. پیشثبت برای علمِ اکتشافی بیشازحد سخت است.
پاسخ. اکتشاف ارزشمند است؛ راهحل برچسبگذاری و جداسازی است. آزادانه کشف کنید، اما فرضیهٔ حاصل را تولیدشده بدانید نه تأییدشده، یا روی دادهٔ تازه اعتبارسنجی کنید.
6با اینها اشتباه نگیرید
مسیرهای منشعب و p-hacking
p-hacking معمولاً جستوجوی هدفمند برای معناداری است؛ مسیرهای منشعب میتواند کاملاً با حسننیت از تصمیمهای پاسخگو به داده ایجاد شود.
چندآزمونی و مسیرهای منشعب
اصلاحِ چندآزمونی خانوادهٔ آزمونهای صریح را میگیرد؛ مسیرهای منشعب آزمونهای ضمنیای را هم شامل میشود که اصلاً اجرا نشدند چون انتخابها به مشاهدهٔ قبلی وابسته بودند.
7خطاهای رایج
- دفاع با «فقط یک مدلِ نهایی اجرا کردیم» وقتی مدلهای جایگزین بهصورتِ ضمنی در دسترس بودند.
- فهمِ پیشثبت بهعنوان ممنوعیتِ اکتشاف بهجای راهی برای جداکردنِ اکتشاف و تأیید.
8در کارِ شما
استدلالی برای ثبتِ پیشاپیشِ طرح، حتی وقتی قصدِ تقلب ندارید. حسنِ نیت دفاع نیست، چون این سازوکار نیازی به سوءنیت ندارد.
9خودآزمایی
اگر پیامد کجتر بود آن را لگاریتمی میکردید، اما چون نبود تحلیلِ خام را گزارش کردید. چرا همین تصمیمِ اجرانشده هم ممکن است استنتاج را متاثر کند؟
نمایش پاسخ
چون قاعدهٔ تحلیل به دادهٔ مشاهدهشده وابسته بود. شاخهٔ خلافواقعی بخشی از روش است و مقدارِ p برای تحلیلِ خامِ ازپیشثابت کلِ فرایندِ انتخاب را توصیف نمیکند.
