PlainLogic

AI explained in plain logic

Why AI hallucinates

Language models predict text, not truth. Why confident wrong answers happen, what causes them, and how to check the answers that matter.

The simple explanation

A language model is trained to predict text, not to guarantee truth. Every response is the most plausible continuation it could produce — and plausible is not the same as verified. Hallucinations come from four directions: missing evidence (it was never shown the fact), misleading prompts (the question smuggles in a false premise), gaps in training (rare topics, new events), and pressure to answer (it always produces something — silence isn’t in the vocabulary). The confident tone is a writing style, not a reliability signal.

It helps to know why training rewards this behavior. The model’s training objective is to predict the next token, and it gets rewarded for fluent, plausible continuations — there’s no separate reward for truth. “I don’t know” is a perfectly good answer for a person, but during training it’s just another string of tokens, competing against thousands of more interesting ones. Later tuning tries to teach honesty, but it’s layered on top of a machine whose core skill is sounding right.

See this idea move.

The Hallucination Lab experiment walks through this concept step by step — press run and watch it happen. Everything is simulated in your browser; no real AI runs.

A concrete example

Ask the Hallucination Lab for the fictional library’s opening year. The source material lists only opening hours. One response invents a year — smoothly and specifically. The other sticks to the evidence and says the year is unknown. Same question, same source — the difference is whether the system is allowed to guess. Always prefer the one that admits what it doesn’t know.

Where people get misled

“It sounded sure, so it’s probably right.” Confidence in AI output is decoration. “It’s usually right, so this one is fine” — a 95% hit rate means nothing for the one answer you needed; errors don’t announce themselves. And “I told it to be careful” — instructions to be careful reduce fabrication but can’t remove it, because the model still can’t check what it doesn’t know.

The practical habit is boring and effective: decide what would prove the answer, then check that. For a date or a name, open the source. For a claim about your own data, ask for the quote and compare it. For anything that moves money or changes a decision, treat the model’s output as a draft that still needs a human signature. Trust is for verified systems, not for fluent ones.

The honest limits

Ground answers in relevant sources and verify the details that matter. Retrieval and instructions to admit uncertainty both help — neither guarantees accuracy. (Our library example is scripted; it demonstrates the concept, it doesn’t measure any real model.)

Plain words on a real concept. The hands-on demo is a simplified educational visualization — it illustrates the idea, not real model internals.