Definition
What is Model calibration?
Model calibration is Model calibration is how closely a model's stated confidence matches how often it is right: a well-calibrated model that says 90% is correct about nine times in ten.
Calibration measures whether a model's confidence can be trusted as a number. If a classifier says "90% likely" about a hundred different cases, a well-calibrated one is right in about ninety of them. If it is right only seventy times, it is overconfident. If it is right ninety-eight times, it is underconfident.
Accuracy and calibration are different
A model can be accurate and badly calibrated. Imagine a spam filter that is right 95% of the time but reports 99.9% confidence on every answer. Its confidence tells you nothing about which answers to double-check. Calibration is the property that makes the number useful for deciding what to do next.
How it is checked
Collect many predictions with their stated confidence, group them by confidence level (all the 60% answers, all the 70% answers and so on), and compare each group's confidence to its actual hit rate. Plotted on a chart, a perfectly calibrated model follows a straight diagonal line. Measures such as expected calibration error summarize the gap in a single number.
Calibration depends on the data. A model can be well calibrated on the kind of cases it was tested on and poorly calibrated on yours, so check it on a sample of your own traffic.
Why it matters to you
Calibrated probabilities let you build a threshold: act automatically above 0.9, send to a person below it. That only works if 0.9 really means about nine out of ten. Without calibration, a threshold is a guess.
This matters most in places where a model makes many small choices: routing support tickets, flagging content, grading an AI agent's actions or picking which tool to call. Cheaper models that return probabilities for a fixed set of options are well suited to these jobs, and their value rests on those probabilities being honest.
Language models writing free text are a harder case. When a chatbot says "I'm certain," that is wording, not a measured probability, and it is often poorly tied to being right. Calibration for text answers is an active research area, and it often needs extra methods such as reading the model's token probabilities.
How it relates to other terms
Calibration concerns the output probabilities of a model, which are built from per-token scores in language models. It is distinct from an AI benchmark, which measures accuracy on a fixed test, and from AI alignment, which concerns goals. A common fix for poor calibration is a light adjustment after training, such as temperature scaling, which rescales the model's scores without changing which answer it picks.
Questions people ask
What does it mean for a model to be well calibrated?
Its stated confidence matches reality: among cases where it says 80%, it is right about 80% of the time.
Is calibration the same as accuracy?
No. Accuracy is how often a model is right. Calibration is whether its confidence number honestly reflects how often it is right.