Why do LLMs hallucinate?
Because a language model is trained to produce plausible text, not to check whether it is true. When the facts it needs are rare in its training data or missing from its context, it still writes a fluent answer, and most benchmarks, and the training tuned to them, reward a guess over saying "I don't know."
Level 1: How AI works · 1.3 Why AI makes things up
AI Evaluation & ReliabilityModel Training
1. The model writes what is plausible, and plausible isn’t true
A large language model produces text by predicting likely next tokens. Nothing in that process checks the result against the world. Dan Klein, co-founder and CTO of Scaled Cognition, put it plainly in episode 54: “If they’re right, we call them truth. And if they’re wrong, we call them hallucination. But the model does not internally distinguish between those.”
So a hallucination isn’t a separate malfunction. It is the same process that produces correct answers, landing on something false. OpenAI’s definition is “plausible but false statements generated by language models.”
2. Some facts can’t be learned from patterns
OpenAI researchers published a paper in September 2025 explaining where these errors come from. In pre-training, a model sees only examples of fluent text, with nothing labeled false. Patterns that repeat, like spelling or matching parentheses, get learned, and errors in them disappear as models scale. Arbitrary facts that appear rarely, like a particular person’s birthday, follow no pattern, so the model can’t predict them from patterns alone. When a question lands on one of those facts, the model still produces something that sounds like an answer.
3. Training and scoreboards reward guessing
The same paper argues that the way models are graded makes this worse. Most benchmarks score accuracy: right answers count, and “I don’t know” scores the same as a wrong answer. A model that guesses a birthday has a 1-in-365 chance of scoring; a model that abstains scores nothing. Over thousands of questions, the guesser looks better. The authors’ fix is to penalize confident errors more than expressions of uncertainty.
Models can decline. Anthropic’s interpretability team found that Claude’s default, when asked about something it doesn’t recognize, is to say it can’t answer. An internal “known entity” signal switches that default off. When the model recognizes a name but knows nothing else about the person, that signal can misfire, and the model answers anyway.
4. Missing context makes it worse
At work, hallucinations often aren’t about trivia. They show up when the model is asked about your refund policy, your customer or your internal acronyms, and that information isn’t in front of it. Chip Huyen, author of AI Engineering, said in episode 5 that “models are more likely to hallucinate when it does not have access to the necessary and correct information to answer questions.” That is why retrieval-augmented generation exists, and why many enterprise hallucinations turn out to be a data problem.
The dangerous ones look right
Obvious errors get caught. Klein’s worry is the ones that don’t. When his team audits enterprise systems, he says, the real hallucination rate is often about five times what the customer has noticed, because the rest are fluent enough to pass. His examples: a refund policy that is a real, plausible policy, but not yours; or yours, but for a customer with a different status. “It looks right and is right are not the same,” he said.
What to do about it
- Decide where a wrong answer is costly. For creative work, invention can be the point. For anything factual, plan for errors.
- Give the model the facts. Put the source documents in its context and ask it to answer only from them, citing where each claim came from.
- Let it say “I don’t know.” Allow it in the instructions, and reward it in your evals instead of scoring it as a miss.
- Check the answers that matter. Test against real questions with known answers, and keep a person on the decisions where a fluent error would do damage.
Hear it from the guest
“If they're right, we call them truth. And if they're wrong, we call them hallucination. But the model does not internally distinguish between those.”
“Looks right. It looks right and is right are not the same. It is very hard to detect hallucinations when they are packaged up sufficiently fluently.”
“Models are more likely to hallucinate when it does not have access to the necessary and correct information to answer questions.”
Quotes lightly edited to remove filler words.
Go deeper
- Why language models hallucinateThe plain-English summary of the paper behind the "training rewards guessing" argument on this page.
- Tracing the thoughts of a large language modelWhat interpretability tools can see inside a model, including the circuits involved in whether it answers or declines.
Common questions
- Will newer models stop hallucinating?
- They hallucinate less, but not to zero. OpenAI says GPT-5 has significantly fewer hallucinations, especially when reasoning, but that they still occur. Its researchers argue accuracy will never reach 100% because some real-world questions are inherently unanswerable. What can improve is how often a model admits it doesn't know instead of guessing.
- Does RAG fix hallucinations?
- It reduces them by putting the right facts in front of the model, which addresses one major cause. It doesn't eliminate them. The model can still misread a source, blend sources, or add claims the sources don't support, so systems that use retrieval need a separate check that each answer is faithful to what was retrieved.
- Can you get a model to say "I don't know"?
- Often, yes. Anthropic's interpretability research found one such mechanism in Claude: its default is to decline when it doesn't recognize a name or entity, and some hallucinations happen when a sense of familiarity wrongly overrides that default. Instructions that allow uncertainty, evals that reward abstaining over guessing, and grounding answers in supplied sources all push in the same direction.
From the conversation
This explainer is drawn from these episodes — each carries its full transcript.
Concepts in this explainer
Large Language ModelPre-trainingGroundingAI HallucinationAI EvaluationAccuracyExplainabilityTokenization