Stage 2 — The Training Data (2023)
Next, the corpus — the thing that was supposed to be the irreducibly human
input — became substantially model-written.
- Phi / phi-1.5 (Microsoft) — Textbooks Are All You Need: a small model
trained on LLM-synthesized, textbook-quality data punched far above its
parameter count.
- WRAP (Apple) — don't just generate data, rephrase the entire web with
an LLM; pretraining gets roughly 3x more efficient.
- Nemotron-4 340B (NVIDIA) — shipped with a permissively licensed
synthetic data generation pipeline as a headline feature.
- Reasoning-trace corpora (2025) — chains of thought generated by strong
reasoners became a standard pretraining and mid-training ingredient.
Synthetic data moved from a trick to an industrial pipeline.