In more detail
Distillation trains a small model to imitate a big one. The large ‘teacher’ model answers vast numbers of questions; the compact ‘student’ is trained on those answers until it echoes the teacher’s judgment — capturing much of the intelligence in a fraction of the size.
It matters because most everyday tasks don’t need a frontier model. Distilled models are how good-enough AI becomes cheap enough to live in your phone keyboard, your email app, your car — the teacher’s wisdom without the teacher’s running costs.
Goes with
