This animated diagram is powered by the PlainLogic flow engine. Follow the packet along the edges — every step is labeled as it fires.
Everything here is a simplified educational visualization — the shape of the idea, not real model internals. The Run button drives a simulated animation in your browser; no real AI model runs and nothing is sent anywhere.
The LLM flow — one token at a time
Simplified educational visualization
A large language model does not look up answers — it writes them one piece at a time. Your prompt is sliced into tokens, the model reads the full context, and predicts the most likely next token, repeating until it stops. Real models do this with probabilities across billions of parameters, so this is a conceptual simplification: no actual tokenization or model math happens here — the animation only shows the shape of the idea.
Ready. Press Run to watch a prompt travel through the flow.
Simplified educational visualization. No real tokenization runs in your browser; the packet animation is illustrative.
Plain-language AI
In plain logic
A language model processes your prompt as tokens — small chunks of text, not whole words. For each position it assigns scores to every token it knows, picks one, appends it to what it is writing, and repeats the whole cycle. The answer grows left to right, one piece at a time.
Training and answering are two different jobs. Training adjusts the model's internal numbers using vast amounts of example text — that is expensive and happens once. Answering (inference) only uses those learned numbers; it does not retrain anything. This is why a model can sound authoritative about something that happened after its training: it cannot learn from your conversation the way you learn from it.
The key insight: the model is a predictor of likely text, not a database of facts. It learned which words tend to follow which other words, and it is very good at continuing patterns. Fluency comes free; truth does not.
Hands-on
Try this
Try completing this sentence in your head: “The cat sat on the ___.” Several endings fit — mat, rug, couch. Your brain picked the most likely one for a generic sentence. A language model does the same thing, over and over, to write whole paragraphs.
Now notice: the most likely ending is not a fact about any real cat. It is just the most common continuation. That gap — between “likely text” and “true statement” — is the entire reason AI hallucinations exist, and it is exactly what the animation above shows: the model picks likely next tokens, nothing more.
Honest boundaries
What this leaves out
Fluent text is not proof. This page shows the shape of generation, but the real machinery — tokenizers, attention, sampling, billions of weights — is deliberately collapsed into six boxes. The simplification is honest but wide.
A model can explain, summarize, or draft with remarkable skill, but important claims still need checking against real sources. Nothing about the animation above measures what any commercial model would actually output for a given prompt — two models, or two settings, can answer the same prompt differently.
Honest answers
Questions people ask
Does a language model understand what it writes?
It tracks patterns well enough to continue them convincingly, but it has no experience of the world behind the words. Treat it as a brilliant pattern-continuation engine, not a thinking mind, and you will calibrate your trust correctly.
Why does the same prompt sometimes get different answers?
The model picks each token from a set of likely candidates, and the pick is sampled with randomness. Change the randomness (the temperature setting) or ask again, and a different-but-plausible continuation can win.
Can I make it smarter by giving it more prompt?
Up to a point. More context helps it pick better tokens — but context has limits, old information gets diluted, and no amount of prompt turns prediction into verified truth. When you need truth, add retrieval (see the RAG guide).
Keep exploring
Related guides
The other seven guides in this series, plus the labs they connect to.