Calibration asks whether events you assign probability p to actually occur about p percent of the time across comparable predictions.
1The problem it solves
Confidence is otherwise hard to audit. Calibration converts subjective probability into a track record: among forecasts called 70%, roughly seven in ten should resolve true. This makes overconfidence and underconfidence measurable and trainable rather than personality labels.
2The idea
A calibrated forecaster is one whose 70% predictions come true about 70% of the time. Calibration is measurable — the Brier score and the logarithmic score both do it — and it is trainable, which is the finding that makes this whole subject practical rather than decorative.
Note that calibration is not the same as accuracy. You can be perfectly calibrated by saying 50% to everything. What you want is calibration plus resolution: confident predictions that turn out right.
3Origin & context
The Brier score, introduced by Glenn Brier, evaluates probabilistic forecasts using squared error; forecasting research associated with Philip Tetlock popularizes calibration training and scoring. Calibration is one dimension of forecast quality, usually paired with resolution or discrimination.
4Canonical example
A forecaster who says 90% and is right 60% of the time is overconfident by a measurable, correctable amount.
5Objections & replies
Objection. A forecaster can be perfectly calibrated by always saying the base rate—50% for everything—while being uninformative.
Reply. Exactly. Calibration is necessary for good probabilistic forecasting but not sufficient. Resolution matters: a good forecaster also distinguishes easy from hard cases and moves probabilities away from the base rate when evidence warrants it.
Objection. Small samples make apparent miscalibration noisy; a few 90% forecasts can easily all fail or all succeed.
Reply. Calibration must be assessed over enough comparable forecasts and with uncertainty about the calibration curve. Individual surprises do not prove miscalibration; systematic deviations over a track record do.
6Don't confuse it with
Calibration and accuracy
Accuracy asks whether individual outcomes were right; calibration asks whether stated probability levels match long-run frequencies. A 60% forecast can be well calibrated even when this particular event fails.
Calibration and discrimination
A calibrated forecaster can still be unhelpfully timid. Discrimination/resolution measures whether forecasts meaningfully separate higher-risk from lower-risk cases.
7Common mistakes
- Judging a probability forecast as 'wrong' solely because the less likely outcome occurred.
- Claiming calibration from a handful of predictions without enough observations at similar probability levels.
8In your work
The highest-return habit available to you: log three to five predictions a week with explicit probabilities and resolution dates, score them quarterly. Nothing else on this map improves your judgment as reliably.
9Check yourself
Across 100 forecasts labeled 80%, only 55 resolve true. What is the primary diagnosis?
Show answer
Systematic overconfidence at that probability level: the stated 80% events occurred about 55% of the time. You would still inspect sample composition and resolution, but the calibration gap is directly measurable.
کالیبراسیون میپرسد آیا رویدادهایی که به آنها احتمالِ p میدهید در مجموعهای از پیشبینیهای مشابه تقریباً p درصدِ مواقع رخ میدهند.
1مسئلهای که حل میکند
در غیر این صورت اطمینان سخت قابلحسابرسی است. کالیبراسیون احتمالِ ذهنی را به کارنامه تبدیل میکند: از پیشبینیهای ۷۰٪ تقریباً هفتدهم باید درست شوند. به این ترتیب بیشاطمینانی و کماطمینانی اندازهگیری و آموزشپذیر میشوند، نه برچسبِ شخصیتی.
2ایدهٔ اصلی
پیشبینیکنندهٔ کالیبره کسی است که پیشبینیهای ۷۰ درصدیاش حدود ۷۰ درصد مواقع درست از آب درمیآید. کالیبراسیون قابل اندازهگیری است — نمرهٔ برایر و نمرهٔ لگاریتمی هر دو این کار را میکنند — و آموختنی است، و همین یافته است که این موضوع را از تزئین به کاربرد بدل میکند.
توجه کنید که کالیبراسیون با دقت یکی نیست. میتوانید با گفتنِ «۵۰ درصد» به همهچیز کاملاً کالیبره باشید. آنچه میخواهید کالیبراسیون بهعلاوهٔ تفکیکپذیری است: پیشبینیهای قاطعی که درست درمیآیند.
3خاستگاه و زمینه
امتیازِ بریر که گلن بریر معرفی کرد پیشبینیِ احتمالی را با خطای مربعی میسنجد؛ پژوهشِ پیشبینی با نامِ فیلیپ تتلاک آموزشِ کالیبراسیون و امتیازدهی را مشهور کرده است. کالیبراسیون فقط یک بُعدِ کیفیتِ پیشبینی است و معمولاً در کنار تفکیک یا قدرتِ تمایز سنجیده میشود.
4مثالِ کلاسیک
پیشبینیکنندهای که ۹۰ درصد میگوید و ۶۰ درصد مواقع درست است، به اندازهای قابل اندازهگیری و قابل اصلاح، بیشاطمینان است.
5اعتراضها و پاسخها
اعتراض. پیشبینیکننده میتواند با گفتنِ همیشگیِ نرخِ پایه — مثلاً ۵۰٪ برای همهچیز — کاملاً کالیبره اما بیفایده باشد.
پاسخ. دقیقاً. کالیبراسیون برای پیشبینیِ احتمالیِ خوب لازم است اما کافی نیست. تفکیک مهم است: پیشبینیکنندهٔ خوب مواردِ آسان و سخت را جدا و وقتی شاهد ایجاب میکند احتمال را از نرخِ پایه دور میکند.
اعتراض. نمونهٔ کوچک کژکالیبراسیونِ ظاهری را پرنویز میکند؛ چند پیشبینیِ ۹۰٪ میتواند اتفاقاً همه شکست بخورد یا همه درست شود.
پاسخ. کالیبراسیون باید روی تعدادِ کافی از پیشبینیهای قابلمقایسه و با عدمِقطعیت دربارهٔ منحنیِ کالیبراسیون سنجیده شود. شگفتیِ منفرد اثباتِ کژکالیبراسیون نیست؛ انحرافِ نظاممند در کارنامه هست.
6با اینها اشتباه نگیرید
کالیبراسیون و دقت
دقت میپرسد نتیجهٔ منفرد درست بود یا نه؛ کالیبراسیون میپرسد سطوحِ احتمالِ اعلامشده با فراوانیِ بلندمدت میخوانند یا نه. پیشبینیِ ۶۰٪ میتواند خوب کالیبره باشد حتی اگر همین مورد شکست بخورد.
کالیبراسیون و تفکیک
پیشبینیکنندهٔ کالیبره میتواند بیش از حد محتاط و کماطلاع باشد. تفکیک میسنجد آیا پیشبینیها واقعاً مواردِ پرخطر را از کمخطر جدا میکنند.
7خطاهای رایج
- «غلط» نامیدنِ پیشبینیِ احتمالی فقط چون نتیجهٔ کماحتمال رخ داد.
- ادعای کالیبراسیون با چند پیشبینی بدون مشاهداتِ کافی در سطوحِ احتمالِ مشابه.
8در کارِ شما
پرثمرترین عادتِ در دسترس: هفتهای سه تا پنج پیشبینی با احتمالِ صریح و تاریخِ تعیینِ نتیجه ثبت کنید و فصلی نمره بدهید. هیچ چیز دیگری در این نقشه به این اطمینان داوریِ شما را بهتر نمیکند.
9خودآزمایی
از ۱۰۰ پیشبینیِ برچسبخورده با ۸۰٪، فقط ۵۵ مورد درست میشوند. تشخیصِ اصلی چیست؟
نمایش پاسخ
بیشاطمینانیِ نظاممند در آن سطح: رویدادهای ۸۰٪ فقط حدود ۵۵٪ رخ دادهاند. هنوز ترکیبِ نمونه و قدرتِ تفکیک را بررسی میکنید، اما شکافِ کالیبراسیون مستقیم اندازهگیری میشود.
