Mastering EpistemologyGuide · Map · Audio فا
An Experiment on a Bird in the Air Pump by Joseph Wright of Derby (1768)

What a p-value isn't

آنچه مقدارِ p نیست

Fisher; the ASA's 2016 statement
Joseph Wright of Derby, An Experiment on a Bird in the Air Pump, 1768

A p-value is a tail-area probability of data under a specified null model; it is not the probability the null is true, the effect is real, or the result will replicate.

1The problem it solves

The most common statistical misreading reverses the conditioning. P(data or more extreme | null) is not P(null | data). A p-value also says nothing by itself about effect size, practical importance, study design, multiple analyses, or prior plausibility. The concept matters because precise language limits what a findings section is entitled to claim.

2The idea

A p-value is the probability of data at least this extreme given that the null hypothesis is true. It is not the probability that the null is true, not the probability the result will replicate, and not a measure of effect size or importance.

3Origin & context

P-values arise from Fisherian significance testing and were later mixed in practice with Neyman–Pearson decision procedures. The American Statistical Association’s 2016 statement emphasized widespread misinterpretations, but the underlying distinction is much older.

4Canonical example

p = 0.04 does not mean a 96% chance the effect is real. It means the data would be unusual if there were no effect.

5Objections & replies

Objection. In common settings, smaller p-values usually correspond to stronger evidence, so the technical warnings are pedantic.

Reply. Often there is a monotonic relation under a fixed model, but evidential strength also depends on alternative hypotheses, stopping rules, multiplicity, power, and prior plausibility. The warnings matter exactly when those conditions differ.

Objection. P-values are so abused they should be banned.

Reply. Abuse is not proof of uselessness. They can summarize incompatibility with a null model under a defined procedure. The solution is to pair them with effect sizes, uncertainty, design, and explicit inferential goals.

6Don't confuse it with

p-value vs. posterior probability

A posterior directly conditions on observed data and requires priors/model alternatives; a p-value conditions on the null and evaluates hypothetical data extremeness.

Statistical significance vs. substantive significance

A tiny effect can be statistically significant with enough data; a large important effect can be imprecisely estimated in a small sample.

7Common mistakes

  • Writing p=0.04 as “there is a 4% chance the null is true.”
  • Treating 0.049 and 0.051 as qualitatively different states of nature.

8In your work

The most consequential misreading in applied work. Fixing it changes what you are entitled to write in a findings section.

9Check yourself

A result has p=0.04. What is the safest short interpretation?

Show answer

Under the specified null model and analysis procedure, data at least this extreme would occur with about 4% probability. It does not by itself give the probability the null is true or the effect will replicate.

مقدارِ p احتمالِ ناحیهٔ دنباله برای داده زیرِ مدلِ صفرِ مشخص است؛ احتمالِ صادق‌بودنِ فرضِ صفر، واقعی‌بودنِ اثر یا تکرارشدنِ نتیجه نیست.

1مسئله‌ای که حل می‌کند

رایج‌ترین بدخوانی شرط را برعکس می‌کند. P(داده یا افراطی‌تر | صفر) برابرِ P(صفر | داده) نیست. p همچنین به‌تنهایی اندازهٔ اثر، اهمیتِ عملی، طراحی، چندتحلیلی یا پیشین را نمی‌گوید. زبانِ دقیق تعیین می‌کند در بخشِ یافته‌ها چه ادعایی مجاز است.

2ایدهٔ اصلی

مقدارِ p احتمالِ دادهٔ دست‌کم به این حد حدی است به‌شرطِ صادق‌بودنِ فرضِ صفر. نه احتمالِ صادق‌بودنِ فرضِ صفر است، نه احتمالِ تکرارِ نتیجه، و نه سنجه‌ای برای اندازه یا اهمیتِ اثر.

3خاستگاه و زمینه

مقدارِ p از آزمونِ معناداریِ فیشر می‌آید و در عمل با روش‌های تصمیمِ نیمن–پیرسن ترکیب شد. بیانیهٔ ۲۰۱۶ انجمنِ آمارِ آمریکا سوءبرداشت‌های گسترده را برجسته کرد، اما تمایزِ منطقی قدیمی‌تر است.

4مثالِ کلاسیک

p = ۰٫۰۴ یعنی ۹۶ درصد احتمالِ واقعی‌بودنِ اثر نیست. یعنی اگر اثری نبود، این داده غیرعادی می‌بود.

5اعتراض‌ها و پاسخ‌ها

اعتراض. در عمل p کوچک‌تر معمولاً شاهدِ قوی‌تر است، پس این هشدارها وسواسی‌اند.

پاسخ. زیرِ مدلِ ثابت اغلب رابطه‌ای هست، اما قوّتِ شاهد به بدیل، توقف، چندگانگی، توان و پیشین هم وابسته است. هشدارها دقیقاً وقتی مهم‌اند که این شرایط فرق کنند.

اعتراض. p آن‌قدر بد استفاده شده که باید حذف شود.

پاسخ. سوءاستفاده بی‌فایدگی را ثابت نمی‌کند. p می‌تواند ناسازگاری با مدلِ صفر را زیرِ رویهٔ معلوم خلاصه کند. باید با اندازهٔ اثر، عدمِ‌قطعیت، طراحی و هدفِ استنتاج همراه شود.

6با این‌ها اشتباه نگیرید

مقدارِ p و احتمالِ پسین

پسین مستقیماً بر داده شرط می‌کند و پیشین/بدیل می‌خواهد؛ p بر فرضِ صفر شرط می‌کند و افراطی‌بودنِ داده‌های فرضی را می‌سنجد.

معناداریِ آماری و اهمیتِ substantive

اثرِ کوچک با دادهٔ زیاد می‌تواند معنادار شود و اثرِ مهم با نمونهٔ کوچک نامطمئن بماند.

7خطاهای رایج

  • نوشتنِ p=0.04 به‌صورتِ «۴٪ احتمال دارد فرضِ صفر درست باشد».
  • رفتار با 0.049 و 0.051 گویی دو وضعیتِ کیفیِ متفاوتِ جهان‌اند.

8در کارِ شما

پرپیامدترین بدخوانی در کارِ کاربردی. اصلاحش تغییر می‌دهد که مجازید چه چیزی در بخشِ یافته‌ها بنویسید.

9خودآزمایی

نتیجه p=0.04 دارد. کوتاه‌ترین تفسیرِ امن چیست؟

نمایش پاسخ

زیرِ مدلِ صفر و رویهٔ تحلیلِ مشخص، داده‌ای به این اندازه یا افراطی‌تر حدود ۴٪ احتمال داشت. این احتمالِ صادق‌بودنِ صفر یا تکرارِ اثر را نمی‌دهد.