The direct answer
Fine-tuning is extra training for an already-trained AI model. You feed it a batch of your own examples — questions paired with the answers you want — and the model's internal settings shift so it reliably behaves the way your examples demonstrate.
Anthropic's glossary puts it plainly: further training on additional data causes the model to "represent and mimic the patterns and characteristics" of that data. Prompting tells the model what to do right now; fine-tuning teaches it a lasting habit.
How it works
You start by collecting examples of the behavior you want, usually as input-and-ideal-output pairs. Then you submit them as a training job through the provider's API — OpenAI's, for instance, accepts a training file and a choice of training method, and returns a new custom model when the job finishes.
During training, the model's weights — the numbers that encode everything it learned in its original training — get nudged, example by example, toward your patterns. The base model's general knowledge stays; the new layer on top is your style, your format, your domain's way of doing things.
A simple example
As an illustration, imagine you run a support desk and want every reply short, calm, and structured the same way. You'd gather a few hundred real tickets paired with ideal replies in that style, run a fine-tuning job on them, and get back a model that answers in that voice by default — no lengthy style instructions needed in every prompt.
Compare that with prompting: there, you'd paste the style guide into every single request, paying for those tokens every time and hoping the model follows them. Fine-tuning moves the instruction from the message into the model.
Why it matters
Fine-tuning earns its keep in three situations. First, consistency at scale: when you need the same format, tone, or behavior thousands of times a day, a fine-tuned model delivers it without a giant prompt. Second, cost and speed: shorter prompts mean fewer tokens per request, and a smaller fine-tuned model can often match a bigger general one on its specialty. Third, behaviors that are hard to describe in words — a brand voice, a house style — which examples convey better than instructions.
OpenAI's own prompt-engineering guide lists prompting techniques first and points to fine-tuning as the later step, which is the right order for most people: prompt until prompting stops being enough.
The common misunderstanding
The common misunderstanding: fine-tuning teaches the model new facts. It's a poor tool for that. Facts baked in during training go stale the moment the world changes, and the model can't tell you which example a fact came from — so there's no source to check.
For knowledge that must stay current or be cited, the right tool is retrieval: fetching documents at request time and handing them to the model alongside the question. Fine-tuning is for patterns — format, tone, style, behavior — not for facts.
What changed recently
Fine-tuning used to mean one thing: show the model more examples. Providers have since widened the toolkit. OpenAI's API now offers three fine-tuning methods — supervised (learn from examples), DPO (learn from preferences: which of two answers is better), and reinforcement (learn from graded outcomes). The idea is the same — keep training the model on your data — but you can now teach judgment, not just imitation.
Try it on PlainLogic
Before reaching for fine-tuning, the Prompt Playbook shows how far good prompting alone can go — most people should start there.
Sources
- PRIMARY SOURCEAnthropic glossary: Fine-tuning
- PRIMARY SOURCEOpenAI API reference: Fine-tuning