AI Glossary

Confidence Calibration

Confidence calibration is how closely predicted probabilities match observed outcomes. Among comparable predictions assigned an 80% chance of being correct, about 80% should be correct if that confidence level is calibrated.

· Updated · Chain of Thought

AI Evaluation & Reliability

Imagine a classifier assigns 80% confidence to 100 ticket categories. If about 80 labels are correct, that group is consistent with calibration at that level. One group is not enough: check across confidence ranges on representative held-out examples, with enough cases to distinguish real patterns from sampling noise.

Calibration differs from accuracy. A model can classify many cases correctly and still overstate its probabilities. The Guo and colleagues paper studies this mismatch in neural classifiers and examines methods for adjusting confidence estimates.

The distinction matters if a product sends uncertain cases to a person or acts on a probability threshold. A language model saying “I am confident” has not supplied a validated correctness probability. Define a measurable confidence signal and an outcome label for the particular task, test their relationship, and recheck it when the workload changes. Calibration does not guarantee that any individual answer is right.

Sources

Go deeper