The direct answer
Inference is the act of running a trained model. You give it new input — a question, a photo, a sentence to translate — and it produces an output using the knowledge baked into its weights. Every chatbot reply, every generated image, every autocomplete suggestion is inference.
Training is the other half of the pair, and the two are opposites in almost every way. Training is a rare, enormously expensive process that adjusts the model's weights using mountains of data. Inference is the cheap, fast, repeated process of using those frozen weights to answer one query at a time. If training is building the engine, inference is driving the car.
How it works
During inference, data flows through the model exactly once — a forward pass — with the weights held fixed. For a chatbot, this happens token by token: each new word is predicted from everything so far, appended, and fed back in. Nothing is learned in the process; the weights that came out of training are byte-for-byte the same after a million queries.
Because inference runs at product scale, there's a whole engineering discipline around making it fast and cheap. Cloud platforms split it into online inference (real-time answers, the model deployed behind an API) and batch inference (scoring piles of data offline). Serving systems squeeze more out of the same GPUs with tricks like continuous batching — packing many users' requests together — memory techniques like PagedAttention, and quantization that shrinks the model so each answer needs less hardware.
A simple example
As an illustration, imagine a vending machine versus the factory that built it. Training is the factory: loud, expensive, and it only has to run once to stamp out the machine. Inference is the vending machine on the corner: quiet, quick, serving one customer after another. And if the machine starts giving wrong change, you fix or replace the machine — you don't rebuild the factory. One chat reply is one press of the button.
Why it matters
Inference is where AI meets economics. Training a frontier model is a one-time cost, however staggering; inference is paid forever, on every single query. That's why API pricing is per token, why companies obsess over smaller and faster models, and why "how much does each answer cost?" is the question that decides whether an AI product survives.
It's also the step that touches your data. Latency, cost per query, and where the servers live are product decisions with real consequences for privacy and user experience — not footnotes. Understanding inference is understanding what you're actually buying when you pay for AI.
The common misunderstanding
The model is not learning from you during inference.
It's natural to feel that a chatbot which "remembers" your conversation is updating itself — but its weights don't change. It simply re-reads your earlier messages from its context window with every reply. End the chat and it knows nothing it didn't know before. Training and inference are separate phases precisely so that using the model can't silently rewrite it. (When a product does personalize to you, that's a separate system layered on top — not inference itself.)
What changed recently
Inference engineering is having a moment. Open-source serving engines like vLLM have made techniques once reserved for big labs — paged memory management, speculative decoding, splitting work across many GPUs — available to anyone running a model. The trend is disaggregation: putting the "read the prompt" phase and the "generate the answer" phase on different hardware, since they stress machines in different ways.
The frontier of AI capability is still training, but the frontier of AI business is increasingly inference. Whoever serves answers cheapest and fastest wins the margin — which is why so much current research is about smaller, sharper, cheaper-to-run models rather than just bigger ones.
Try it on PlainLogic
The AI Lab on PlainLogic shows models doing exactly what this article describes: loaded weights, live input, answers in real time — inference you can watch happen.
Sources
- PRIMARY SOURCEGoogle Cloud: Model training and inference workflow (Vertex AI docs)
- PRIMARY SOURCEvLLM: LLM inference and serving documentation