PlainLogic

PlainLogic Explainer

What is inference? Running a trained model vs. training it

Inference is simply using a trained AI model: new data goes in, a prediction comes out. Training builds the model once; inference runs it millions of times. Almost everything you pay for in AI is inference.

The direct answer

Inference is the act of running a trained model. You give it new input — a question, a photo, a sentence to translate — and it produces an output using the knowledge baked into its weights. Every chatbot reply, every generated image, every autocomplete suggestion is inference.

Training is the other half of the pair, and the two are opposites in almost every way. Training is a rare, enormously expensive process that adjusts the model's weights using mountains of data. Inference is the cheap, fast, repeated process of using those frozen weights to answer one query at a time. If training is building the engine, inference is driving the car.

How it works

During inference, data flows through the model exactly once — a forward pass — with the weights held fixed. For a chatbot, this happens token by token: each new word is predicted from everything so far, appended, and fed back in. Nothing is learned in the process; the weights that came out of training are byte-for-byte the same after a million queries.

Because inference runs at product scale, there's a whole engineering discipline around making it fast and cheap. Cloud platforms split it into online inference (real-time answers, the model deployed behind an API) and batch inference (scoring piles of data offline). Serving systems squeeze more out of the same GPUs with tricks like continuous batching — packing many users' requests together — memory techniques like PagedAttention, and quantization that shrinks the model so each answer needs less hardware.

A simple example

As an illustration, imagine a vending machine versus the factory that built it. Training is the factory: loud, expensive, and it only has to run once to stamp out the machine. Inference is the vending machine on the corner: quiet, quick, serving one customer after another. And if the machine starts giving wrong change, you fix or replace the machine — you don't rebuild the factory. One chat reply is one press of the button.

Why it matters

Inference is where AI meets economics. Training a frontier model is a one-time cost, however staggering; inference is paid forever, on every single query. That's why API pricing is per token, why companies obsess over smaller and faster models, and why "how much does each answer cost?" is the question that decides whether an AI product survives.

It's also the step that touches your data. Latency, cost per query, and where the servers live are product decisions with real consequences for privacy and user experience — not footnotes. Understanding inference is understanding what you're actually buying when you pay for AI.

The common misunderstanding

The model is not learning from you during inference.

It's natural to feel that a chatbot which "remembers" your conversation is updating itself — but its weights don't change. It simply re-reads your earlier messages from its context window with every reply. End the chat and it knows nothing it didn't know before. Training and inference are separate phases precisely so that using the model can't silently rewrite it. (When a product does personalize to you, that's a separate system layered on top — not inference itself.)

What changed recently

Inference engineering is having a moment. Open-source serving engines like vLLM have made techniques once reserved for big labs — paged memory management, speculative decoding, splitting work across many GPUs — available to anyone running a model. The trend is disaggregation: putting the "read the prompt" phase and the "generate the answer" phase on different hardware, since they stress machines in different ways.

The frontier of AI capability is still training, but the frontier of AI business is increasingly inference. Whoever serves answers cheapest and fastest wins the margin — which is why so much current research is about smaller, sharper, cheaper-to-run models rather than just bigger ones.

Try it on PlainLogic

The AI Lab on PlainLogic shows models doing exactly what this article describes: loaded weights, live input, answers in real time — inference you can watch happen.

Sources

How this was made: PlainLogic uses automation to monitor technology updates and assist with research and drafting. Articles are built from cited sources and checked for factual consistency before publication.

Keep learning

Keep learning

More plain-words guides and hands-on experiments from the PlainLogic lab.