PlainLogic

PlainLogic Explainer

What is model distillation? Teaching small AI models to think big

Model distillation is how a small, fast AI model learns from a big, slow one. Instead of studying raw data alone, the student model trains on the teacher's judgment — and ends up far smarter than its size suggests.

The direct answer

Model distillation is a training technique where a small neural network (the “student”) learns from a large, already-trained one (the “teacher”). Instead of training the small model only on labeled data — right answers — you also train it to mimic the teacher's outputs on lots of examples. The small model ends up performing much closer to the big one than it ever could on its own, while staying fast and cheap to run. The term was coined by Geoffrey Hinton, Oriol Vinyals, and Jeff Dean in their 2015 paper “Distilling the Knowledge in a Neural Network.”

The idea is older than the name. In 2006, Bucilă, Caruana, and Niculescu-Mizil showed that a whole ensemble of models could be compressed into one smaller model without much accuracy loss. Distillation formalized the trick and made it standard practice — it's one reason so much AI now runs on your phone instead of a distant server farm.

How it works

The key insight is that a teacher's mistakes are informative. When a big image classifier looks at a photo of a dog, it doesn't just say “dog” — it quietly reports a whole ranking: very likely a dog, somewhat likely a wolf, barely likely a cat. That ranking is called a soft target, and it reveals how the teacher generalizes: dogs and wolves share features, dogs and cars don't. A student trained to reproduce the full ranking absorbs the teacher's sense of similarity between classes, not just the final labels.

In practice, you “soften” the teacher's outputs with a temperature setting that spreads out its probabilities, making those faint runner-up signals visible; the student trains against them. After training, you turn the temperature back down and the small model answers with normal confidence. The student never needs the teacher's size or its training data — just its behavior on a pile of examples, labeled or not.

A simple example

An illustration.

Imagine a master chef training an apprentice. The apprentice could learn by tasting finished dishes and guessing the recipe — that's ordinary training on labeled data. But it's far faster to watch the chef cook: seeing how much salt she adds, when she tastes, what she fixes mid-dish. Distillation is the apprentice watching the chef, not just tasting the dish — the student model gets the teacher's process, expressed through its soft, uncertain judgments.

Why it matters

Size is the bottleneck of modern AI. The best models are enormous — expensive to run, slow to answer, and hungry for hardware. Distillation breaks that tradeoff: you train one giant model once, then distill it into a family of smaller models tuned for different jobs. That small model can then run on a laptop, a phone, or inside a product where every millisecond and every cent of compute matters.

It also changes what “small” means. A distilled small model routinely beats a small model trained from scratch on the same data, because the teacher's soft targets carry information the raw labels don't — the researchers call it “dark knowledge.” For anyone building AI products, distillation is the reason capable models can live inside apps rather than behind expensive API calls.

The common misunderstanding

The common misunderstanding is that the student just memorizes the teacher's answers — a copy, only dumber. What it actually copies is the teacher's uncertainty: which classes the teacher considers near-misses, how confident it is, where it hedges. Two students can agree with the teacher on every final answer and still behave differently, because one learned the teacher's judgment and the other only its verdicts. That's also why a distilled student can occasionally beat its teacher on specific tasks: the teacher's soft targets act like a gentle regularizer, smoothing out the teacher's own overconfident mistakes.

What changed recently

Distillation has gone from a compression trick to a production strategy. Many of today's compact models are distilled from much larger ones, and the technique keeps evolving — self-distillation, where a model teaches a copy of itself to squeeze out extra performance, and multi-teacher setups that blend several big models into one student. The frontier question is how far down you can go: how small a model can still carry the knowledge of a giant, and at what point the student's capacity simply runs out of room.

Try it on PlainLogic

The AI Lab explores how models of different sizes get built and run — the world distillation makes possible.

Sources

How this was made: PlainLogic uses automation to monitor technology updates and assist with research and drafting. Articles are built from cited sources and checked for factual consistency before publication.