Stage 5 — The Researcher (2026)

The assistance era (Copilot, then SWE-agents) kept a human choosing the

experiments. The discovery era did not.

later published in Nature.

minimal ratchet loop where a coding agent modifies a real LLM training

setup, runs a five-minute experiment, keeps the change only if validation

loss improves, and repeats overnight. His extended run stacked 700

experiments into 20 kept improvements, cutting time-to-GPT-2 from 2.02 to

1.80 hours — real, transferable code changes found while he slept.

The pattern replicated within weeks, turning one GitHub repo into a genre:

experiments, a 2.3% validation-loss improvement, zero human intervention.

ran 333 experiments in a single night (March 8–9, 2026).

workload.

mid-2026 the loop itself had become something you score agents on.

Experiment selection — the last intellectual step humans kept in the loop —

became an optimization target, then a benchmark category.