AI, decoded

What is model distillation in AI?

Model distillation trains a student model to reproduce useful behavior from a teacher model, often transferring a narrow capability into a smaller model. The training signal may be the teacher’s answers or its predicted probabilities; the student still needs independent testing on the task it will perform.

· Chain of Thought

Level 4: How AI systems are built · 4.5 Customizing models

Model TrainingModel Architecture

A teacher supplies the learning signal

In episode 43, Maxime Labonne describes a teacher model passing knowledge to a student during training. Alex Ratner makes the same distinction in episode 57: generating data with a stronger model can have a useful purpose when that data trains a smaller one. It is a transfer process with a source of supervision.

The distillation paper by Hinton, Vinyals and Dean describes training a smaller model using the outputs of a larger model or ensemble. Predicted probabilities can convey more information than a single winning label: they express which alternative answers the teacher considers plausible. That is one form of distillation, rather than a requirement for every workflow bearing the name.

Decide which behavior to transfer

Start with a bounded task and a clear test of success. Illustrative example: an expensive teacher classifies incoming messages into billing, delivery and account-access queues. You want a student to do that same routing. You are not trying to reproduce everything the teacher knows or every kind of conversation it can hold.

Build examples across the queues, including messages with several requests and messages that do not fit any queue. Have a reviewer check the teacher’s labels against the routing policy. Keep a separate evaluation set with independently established answers, then compare teacher and student on it.

Record where the student disagrees. A disagreement is something to inspect, not automatic proof that the student is wrong. The teacher may have mislabeled the message. For a difficult case, retain the evidence and the policy decision that settled the label.

Distillation and fine-tuning can overlap

The terms describe different parts of the process. Distillation identifies the teacher as the source of the learning signal. Supervised fine-tuning describes training on examples of the desired response. A workflow that fine-tunes a student on reviewed teacher answers can be both.

Other methods compare teacher and student predictions during training. Hugging Face’s distillation documentation provides an implementation for that kind of teacher-student training. You do not need those implementation details to make the first product decision: which task is worth transferring, and what would count as a successful transfer?

For the routing example, compare correct destinations, mistaken escalations, missed escalations and response time. Include the cost of preparing data and operating the student when deciding whether the change is worthwhile.

Where it falls short

A student can inherit a teacher’s mistakes, and a narrow training set can leave important situations uncovered. Ratner warns against treating generated data as a machine that creates missing expertise by itself. Teacher output is evidence to review, not an independent answer key.

Do not assume a student that routes messages well can also answer them. Keep its responsibility explicit, with a fallback for unsupported cases. Continue with preparing fine-tuning data to work through coverage and review before training.

Hear it from the guest

“Having a teacher model that is able to distill some knowledge into the student model that you're currently training.”
“One of those is called distillation, where you're using a bigger model to train a smaller model.”

Quotes lightly edited to remove filler words.

Go deeper

From the conversation

This explainer is drawn from these episodes — each carries its full transcript.

Concepts in this explainer

Knowledge DistillationFine-TuningSupervised Fine-TuningSynthetic Data