Stage 3 — The Teacher (2023)
Weeks after ChatGPT's API opened, imitation learning industrialized — the
teacher is a model.
- Alpaca (Stanford) — a $600 fine-tune on GPT-generated instructions
cloned much of a frontier model's behavior.
- Vicuna — trained on shared user conversations scraped from the web.
- Orca (Microsoft) — rich teacher explanations rather than bare answers.
- On-policy distillation — matured imitation into a proper training
discipline, fixing the train/inference mismatch.
- DeepSeek-R1 — shipped a family of distilled models alongside the
flagship, making "the teacher is a model" the default assumption for
small-model releases since.
Education — explanation, demonstration, correction — turned out to be
perfectly simulable.