Mastering EpistemologyGuide · Map · Audio فا
Impression, Sunrise by Claude Monet (1872)

Calibration and scoring

کالیبراسیون و نمره‌دهی

Brier · Tetlock
Claude Monet, Impression, Sunrise, 1872

Calibration asks whether events you assign probability p to actually occur about p percent of the time across comparable predictions.

1The problem it solves

Confidence is otherwise hard to audit. Calibration converts subjective probability into a track record: among forecasts called 70%, roughly seven in ten should resolve true. This makes overconfidence and underconfidence measurable and trainable rather than personality labels.

2The idea

A calibrated forecaster is one whose 70% predictions come true about 70% of the time. Calibration is measurable — the Brier score and the logarithmic score both do it — and it is trainable, which is the finding that makes this whole subject practical rather than decorative.

Note that calibration is not the same as accuracy. You can be perfectly calibrated by saying 50% to everything. What you want is calibration plus resolution: confident predictions that turn out right.

3Origin & context

The Brier score, introduced by Glenn Brier, evaluates probabilistic forecasts using squared error; forecasting research associated with Philip Tetlock popularizes calibration training and scoring. Calibration is one dimension of forecast quality, usually paired with resolution or discrimination.

4Canonical example

A forecaster who says 90% and is right 60% of the time is overconfident by a measurable, correctable amount.

5Objections & replies

Objection. A forecaster can be perfectly calibrated by always saying the base rate—50% for everything—while being uninformative.

Reply. Exactly. Calibration is necessary for good probabilistic forecasting but not sufficient. Resolution matters: a good forecaster also distinguishes easy from hard cases and moves probabilities away from the base rate when evidence warrants it.

Objection. Small samples make apparent miscalibration noisy; a few 90% forecasts can easily all fail or all succeed.

Reply. Calibration must be assessed over enough comparable forecasts and with uncertainty about the calibration curve. Individual surprises do not prove miscalibration; systematic deviations over a track record do.

6Don't confuse it with

Calibration and accuracy

Accuracy asks whether individual outcomes were right; calibration asks whether stated probability levels match long-run frequencies. A 60% forecast can be well calibrated even when this particular event fails.

Calibration and discrimination

A calibrated forecaster can still be unhelpfully timid. Discrimination/resolution measures whether forecasts meaningfully separate higher-risk from lower-risk cases.

7Common mistakes

  • Judging a probability forecast as 'wrong' solely because the less likely outcome occurred.
  • Claiming calibration from a handful of predictions without enough observations at similar probability levels.

8In your work

The highest-return habit available to you: log three to five predictions a week with explicit probabilities and resolution dates, score them quarterly. Nothing else on this map improves your judgment as reliably.

9Check yourself

Across 100 forecasts labeled 80%, only 55 resolve true. What is the primary diagnosis?

Show answer

Systematic overconfidence at that probability level: the stated 80% events occurred about 55% of the time. You would still inspect sample composition and resolution, but the calibration gap is directly measurable.

کالیبراسیون می‌پرسد آیا رویدادهایی که به آن‌ها احتمالِ p می‌دهید در مجموعه‌ای از پیش‌بینی‌های مشابه تقریباً p درصدِ مواقع رخ می‌دهند.

1مسئله‌ای که حل می‌کند

در غیر این صورت اطمینان سخت قابل‌حسابرسی است. کالیبراسیون احتمالِ ذهنی را به کارنامه تبدیل می‌کند: از پیش‌بینی‌های ۷۰٪ تقریباً هفت‌دهم باید درست شوند. به این ترتیب بیش‌اطمینانی و کم‌اطمینانی اندازه‌گیری و آموزش‌پذیر می‌شوند، نه برچسبِ شخصیتی.

2ایدهٔ اصلی

پیش‌بینی‌کنندهٔ کالیبره کسی است که پیش‌بینی‌های ۷۰ درصدی‌اش حدود ۷۰ درصد مواقع درست از آب درمی‌آید. کالیبراسیون قابل اندازه‌گیری است — نمرهٔ برایر و نمرهٔ لگاریتمی هر دو این کار را می‌کنند — و آموختنی است، و همین یافته است که این موضوع را از تزئین به کاربرد بدل می‌کند.

توجه کنید که کالیبراسیون با دقت یکی نیست. می‌توانید با گفتنِ «۵۰ درصد» به همه‌چیز کاملاً کالیبره باشید. آنچه می‌خواهید کالیبراسیون به‌علاوهٔ تفکیک‌پذیری است: پیش‌بینی‌های قاطعی که درست درمی‌آیند.

3خاستگاه و زمینه

امتیازِ بریر که گلن بریر معرفی کرد پیش‌بینیِ احتمالی را با خطای مربعی می‌سنجد؛ پژوهشِ پیش‌بینی با نامِ فیلیپ تتلاک آموزشِ کالیبراسیون و امتیازدهی را مشهور کرده است. کالیبراسیون فقط یک بُعدِ کیفیتِ پیش‌بینی است و معمولاً در کنار تفکیک یا قدرتِ تمایز سنجیده می‌شود.

4مثالِ کلاسیک

پیش‌بینی‌کننده‌ای که ۹۰ درصد می‌گوید و ۶۰ درصد مواقع درست است، به اندازه‌ای قابل اندازه‌گیری و قابل اصلاح، بیش‌اطمینان است.

5اعتراض‌ها و پاسخ‌ها

اعتراض. پیش‌بینی‌کننده می‌تواند با گفتنِ همیشگیِ نرخِ پایه — مثلاً ۵۰٪ برای همه‌چیز — کاملاً کالیبره اما بی‌فایده باشد.

پاسخ. دقیقاً. کالیبراسیون برای پیش‌بینیِ احتمالیِ خوب لازم است اما کافی نیست. تفکیک مهم است: پیش‌بینی‌کنندهٔ خوب مواردِ آسان و سخت را جدا و وقتی شاهد ایجاب می‌کند احتمال را از نرخِ پایه دور می‌کند.

اعتراض. نمونهٔ کوچک کژکالیبراسیونِ ظاهری را پرنویز می‌کند؛ چند پیش‌بینیِ ۹۰٪ می‌تواند اتفاقاً همه شکست بخورد یا همه درست شود.

پاسخ. کالیبراسیون باید روی تعدادِ کافی از پیش‌بینی‌های قابل‌مقایسه و با عدمِ‌قطعیت دربارهٔ منحنیِ کالیبراسیون سنجیده شود. شگفتیِ منفرد اثباتِ کژکالیبراسیون نیست؛ انحرافِ نظام‌مند در کارنامه هست.

6با این‌ها اشتباه نگیرید

کالیبراسیون و دقت

دقت می‌پرسد نتیجهٔ منفرد درست بود یا نه؛ کالیبراسیون می‌پرسد سطوحِ احتمالِ اعلام‌شده با فراوانیِ بلندمدت می‌خوانند یا نه. پیش‌بینیِ ۶۰٪ می‌تواند خوب کالیبره باشد حتی اگر همین مورد شکست بخورد.

کالیبراسیون و تفکیک

پیش‌بینی‌کنندهٔ کالیبره می‌تواند بیش از حد محتاط و کم‌اطلاع باشد. تفکیک می‌سنجد آیا پیش‌بینی‌ها واقعاً مواردِ پرخطر را از کم‌خطر جدا می‌کنند.

7خطاهای رایج

  • «غلط» نامیدنِ پیش‌بینیِ احتمالی فقط چون نتیجهٔ کم‌احتمال رخ داد.
  • ادعای کالیبراسیون با چند پیش‌بینی بدون مشاهداتِ کافی در سطوحِ احتمالِ مشابه.

8در کارِ شما

پرثمرترین عادتِ در دسترس: هفته‌ای سه تا پنج پیش‌بینی با احتمالِ صریح و تاریخِ تعیینِ نتیجه ثبت کنید و فصلی نمره بدهید. هیچ چیز دیگری در این نقشه به این اطمینان داوریِ شما را بهتر نمی‌کند.

9خودآزمایی

از ۱۰۰ پیش‌بینیِ برچسب‌خورده با ۸۰٪، فقط ۵۵ مورد درست می‌شوند. تشخیصِ اصلی چیست؟

نمایش پاسخ

بیش‌اطمینانیِ نظام‌مند در آن سطح: رویدادهای ۸۰٪ فقط حدود ۵۵٪ رخ داده‌اند. هنوز ترکیبِ نمونه و قدرتِ تفکیک را بررسی می‌کنید، اما شکافِ کالیبراسیون مستقیم اندازه‌گیری می‌شود.