Stage 2 — The Training Data (2023)

Next, the corpus — the thing that was supposed to be the irreducibly human

input — became substantially model-written.

trained on LLM-synthesized, textbook-quality data punched far above its

parameter count.

an LLM; pretraining gets roughly 3x more efficient.

synthetic data generation pipeline as a headline feature.

reasoners became a standard pretraining and mid-training ingredient.

Synthetic data moved from a trick to an industrial pipeline.