A p-value is a tail-area probability of data under a specified null model; it is not the probability the null is true, the effect is real, or the result will replicate.
1The problem it solves
The most common statistical misreading reverses the conditioning. P(data or more extreme | null) is not P(null | data). A p-value also says nothing by itself about effect size, practical importance, study design, multiple analyses, or prior plausibility. The concept matters because precise language limits what a findings section is entitled to claim.
2The idea
A p-value is the probability of data at least this extreme given that the null hypothesis is true. It is not the probability that the null is true, not the probability the result will replicate, and not a measure of effect size or importance.
3Origin & context
P-values arise from Fisherian significance testing and were later mixed in practice with Neyman–Pearson decision procedures. The American Statistical Association’s 2016 statement emphasized widespread misinterpretations, but the underlying distinction is much older.
4Canonical example
p = 0.04 does not mean a 96% chance the effect is real. It means the data would be unusual if there were no effect.
5Objections & replies
Objection. In common settings, smaller p-values usually correspond to stronger evidence, so the technical warnings are pedantic.
Reply. Often there is a monotonic relation under a fixed model, but evidential strength also depends on alternative hypotheses, stopping rules, multiplicity, power, and prior plausibility. The warnings matter exactly when those conditions differ.
Objection. P-values are so abused they should be banned.
Reply. Abuse is not proof of uselessness. They can summarize incompatibility with a null model under a defined procedure. The solution is to pair them with effect sizes, uncertainty, design, and explicit inferential goals.
6Don't confuse it with
p-value vs. posterior probability
A posterior directly conditions on observed data and requires priors/model alternatives; a p-value conditions on the null and evaluates hypothetical data extremeness.
Statistical significance vs. substantive significance
A tiny effect can be statistically significant with enough data; a large important effect can be imprecisely estimated in a small sample.
7Common mistakes
- Writing p=0.04 as “there is a 4% chance the null is true.”
- Treating 0.049 and 0.051 as qualitatively different states of nature.
8In your work
The most consequential misreading in applied work. Fixing it changes what you are entitled to write in a findings section.
9Check yourself
A result has p=0.04. What is the safest short interpretation?
Show answer
Under the specified null model and analysis procedure, data at least this extreme would occur with about 4% probability. It does not by itself give the probability the null is true or the effect will replicate.
مقدارِ p احتمالِ ناحیهٔ دنباله برای داده زیرِ مدلِ صفرِ مشخص است؛ احتمالِ صادقبودنِ فرضِ صفر، واقعیبودنِ اثر یا تکرارشدنِ نتیجه نیست.
1مسئلهای که حل میکند
رایجترین بدخوانی شرط را برعکس میکند. P(داده یا افراطیتر | صفر) برابرِ P(صفر | داده) نیست. p همچنین بهتنهایی اندازهٔ اثر، اهمیتِ عملی، طراحی، چندتحلیلی یا پیشین را نمیگوید. زبانِ دقیق تعیین میکند در بخشِ یافتهها چه ادعایی مجاز است.
2ایدهٔ اصلی
مقدارِ p احتمالِ دادهٔ دستکم به این حد حدی است بهشرطِ صادقبودنِ فرضِ صفر. نه احتمالِ صادقبودنِ فرضِ صفر است، نه احتمالِ تکرارِ نتیجه، و نه سنجهای برای اندازه یا اهمیتِ اثر.
3خاستگاه و زمینه
مقدارِ p از آزمونِ معناداریِ فیشر میآید و در عمل با روشهای تصمیمِ نیمن–پیرسن ترکیب شد. بیانیهٔ ۲۰۱۶ انجمنِ آمارِ آمریکا سوءبرداشتهای گسترده را برجسته کرد، اما تمایزِ منطقی قدیمیتر است.
4مثالِ کلاسیک
p = ۰٫۰۴ یعنی ۹۶ درصد احتمالِ واقعیبودنِ اثر نیست. یعنی اگر اثری نبود، این داده غیرعادی میبود.
5اعتراضها و پاسخها
اعتراض. در عمل p کوچکتر معمولاً شاهدِ قویتر است، پس این هشدارها وسواسیاند.
پاسخ. زیرِ مدلِ ثابت اغلب رابطهای هست، اما قوّتِ شاهد به بدیل، توقف، چندگانگی، توان و پیشین هم وابسته است. هشدارها دقیقاً وقتی مهماند که این شرایط فرق کنند.
اعتراض. p آنقدر بد استفاده شده که باید حذف شود.
پاسخ. سوءاستفاده بیفایدگی را ثابت نمیکند. p میتواند ناسازگاری با مدلِ صفر را زیرِ رویهٔ معلوم خلاصه کند. باید با اندازهٔ اثر، عدمِقطعیت، طراحی و هدفِ استنتاج همراه شود.
6با اینها اشتباه نگیرید
مقدارِ p و احتمالِ پسین
پسین مستقیماً بر داده شرط میکند و پیشین/بدیل میخواهد؛ p بر فرضِ صفر شرط میکند و افراطیبودنِ دادههای فرضی را میسنجد.
معناداریِ آماری و اهمیتِ substantive
اثرِ کوچک با دادهٔ زیاد میتواند معنادار شود و اثرِ مهم با نمونهٔ کوچک نامطمئن بماند.
7خطاهای رایج
- نوشتنِ p=0.04 بهصورتِ «۴٪ احتمال دارد فرضِ صفر درست باشد».
- رفتار با 0.049 و 0.051 گویی دو وضعیتِ کیفیِ متفاوتِ جهاناند.
8در کارِ شما
پرپیامدترین بدخوانی در کارِ کاربردی. اصلاحش تغییر میدهد که مجازید چه چیزی در بخشِ یافتهها بنویسید.
9خودآزمایی
نتیجه p=0.04 دارد. کوتاهترین تفسیرِ امن چیست؟
نمایش پاسخ
زیرِ مدلِ صفر و رویهٔ تحلیلِ مشخص، دادهای به این اندازه یا افراطیتر حدود ۴٪ احتمال داشت. این احتمالِ صادقبودنِ صفر یا تکرارِ اثر را نمیدهد.
