Stage 6 — The Environment (2026)

RL's scaling bottleneck moved from the model to the environment: you need

thousands of executable, verifiable, professionally realistic task worlds,

and humans can't hand-build them fast enough.

research agents mine real work patterns and convert them into long-horizon

environments with hidden state; a judge agent attempts each task to

confirm it's solvable; verifiers are synthesized *without seeing the

reference solution*, then stress-tested with oracle, no-op, and

unsolved-state checks until their binary reward is reliable enough to

train on directly. The environment, judging, and verification stack is

synthetic all the way down.

its own tasks and generates its own RL rollouts.

By mid-2026 the GLM-5.3 recipe had been productized and open-sourced:

SQL-backed executable environments for agentic RL with benchmark-winning

results; environment synthesis is now an arXiv subfield in its own right.

build → judge-agent solvability → verify → train) built on agent SDKs such

as Claude's, with a frontier model as the backbone of every module.

to anchor a summer school (FoPSS 2026), covering generate–verify–retry

loops and binary-reward reliability.

Late 2026 pushed the stage from "synthesize a gym" to "co-evolve with the

agent":

policy (WebEvolver, DreamGym): the environment adapts to the agent while

the agent learns, instead of staying fixed.

two-player game: a Synthesizer agent invents tasks, a Solver attempts

them, and the game itself generates the curriculum.

(e.g. RL List, Oct 2026) and funded startups selling simulation sandboxes

for benchmarking and training agents before deployment. Environment

synthesis is a procurement category now, not just a paper topic.

The gym, the referee, and the scoreboard are all models now — the referee

is getting its own theory, and the gym has started fighting back.