---
id: PRG-0092
title: I Do Not Know Is The Answer Nobody Grades
kicker: the assurance desk, on calibration
captured: 2026-09-30T14:30:00Z
status: open
author: Marlowe Quist
summary: A language model that invents an answer is behaving exactly as it was scored to behave, because most evaluations give zero credit for abstaining. The upside of AI depends on a single measurable property, calibration, and the way models are graded quietly trains it away.
tags: [assurance, calibration, the record, permission, ai]
---

Calibration is the agreement between how sure a system says it is and how often it is right. A forecaster who announces seventy percent chance of rain is calibrated if, across every day she said seventy, it rained on about seven in ten. The property says nothing about whether she is a good forecaster. It says whether her confidence can be used as a number. For a deployed AI system that is the more important question, because a confidence you can use is the thing that lets a human decide when to check.

On the ledger I keep between people and tokens, calibration is the line that converts a model's raw accuracy into actual value. A model that is right ninety percent of the time and cannot tell you which ninety percent forces a person to check everything, which means the saved labor was never saved. A model that is right ninety percent of the time and reliably flags the other ten lets the person check only the flags. <Highlight>The second model is worth many times the first, and the leaderboard records them as identical.</Highlight>

## how the guessing gets trained in

Picture a multiple-choice exam with no penalty for wrong answers. A student who leaves a question blank scores zero on it. A student who guesses scores zero or one. Over a long exam the guesser always wins, so every sensible student learns to guess. Most benchmarks used to rank language models are scored this way: right or wrong, with no partial credit for saying the question cannot be answered from what the model knows.

A 2025 paper from researchers at OpenAI and Georgia Tech, [Why Language Models Hallucinate](https://arxiv.org/abs/2509.04664), makes this argument formally. Some fabrication falls out of pretraining statistics, they show, because facts that appear only once in the training data are very hard to learn reliably. The persistence of fabrication after all the subsequent tuning, they argue, is an incentive problem. Evaluations reward a confident guess over an honest abstention, so the training pipeline that optimizes against those evaluations produces confident guessers.

There is earlier evidence pointing the same direction. In 2017 a group at Cornell showed that [modern deep neural networks are systematically overconfident](https://arxiv.org/abs/1706.04599) compared with older, smaller ones, and that a one-parameter adjustment called temperature scaling recovers most of the lost calibration. In 2023 OpenAI's GPT-4 technical report published a striking chart: the base model, before tuning for helpfulness, was closely calibrated on a multiple-choice test, and the tuned assistant was noticeably less so. The process that made it pleasant to talk to also made its confidence less informative.

> A model that never says it does not know has been taught that its silence is worth nothing.

## the reject option is older than the field

None of this is a new idea in statistics. In 1970 C. K. Chow, an engineer working on character recognition, described the optimal tradeoff between error and rejection: a classifier allowed to decline the hardest cases can drive its error rate on the cases it accepts sharply down. Every serious deployment since has rediscovered the same structure under different names. Selective prediction. Learning to defer. Human in the loop.

What changes the value of a deployment is the routing:

- **Accepted and right.** Labor actually saved. This is the line on the slide.
- **Rejected and routed.** A person handles a hard case, with the model having narrowed the search. Still a gain, smaller and more honest.
- **Accepted and wrong, flagged low-confidence.** Recoverable, if a reviewer is reading flags.
- **Accepted and wrong, stated with confidence.** The cost that leaves the building. It is paid later, by someone the vendor never meets.

That last row is where hallucination does its damage, and calibration is the only property that shrinks it without shrinking the first row along with it.

<Marginalia label="On the human sensor">Every reviewer in a selective-prediction loop is doing something the model cannot do for itself: noticing that a confident answer smells wrong. I treat those people as instruments. When their flags start clustering on one category of question, that clustering is the first signal that the model's calibration has drifted, usually weeks before any dashboard shows it. Remove the reviewers to save money and you remove the sensor that told you the money was safe.</Marginalia>

## the position

The upside here is real and specific. In fields where errors are expensive and checks are cheap, a calibrated model with an abstain path compresses expert time without hiding risk, and that is most of what anyone actually wants from these systems. The downside is equally specific: an uncalibrated model trained to guess will look identical on a leaderboard and produce its worst errors with its steadiest voice.

So I hold every deployment to three checks before I will sign. Does the system expose a confidence that has been measured against outcomes, on data from the deployment and not the vendor's demo? Is there a path where the model is allowed to decline, and is declining scored as a success when it is correct? And is someone accountable for reading what gets declined? The [trust a system earns should be adaptive](https://www.adjective.us/blog/agent-trust-plane-earned-autonomy), widening as its confidence proves out and narrowing when it drifts, rather than granted once in a policy file. The authorization to act on a confident answer is itself a record, and it should be [written and sealed with the evidence](https://www.adjective.us/evidence-sealed-authorization) that justified it, so that later, when a confident answer turns out wrong, there is a document showing who decided that confidence was enough.

[Assurance is an engineering discipline](https://www.adjective.us/blog/ai-assurance-is-not-a-policy-problem), and calibration is one of the few parts of it you can actually measure. Measure it.

Grade the abstention, or the machine will learn that pretending is the job.
