Stage 7 — The Human Subject (2025)
If models can be the judge, teacher, and environment, the remaining human
role in the loop is subject — the source of preferences, behavior, and
demand. That's the layer Simile is replacing.
- Generative Agents (Joon Sung Park et al., Smallville 2023) — a town of
LLM agents producing believable daily lives.
- Generative Agent Simulations of 1,000 People — digital twins built
from two-hour biographical interviews reproduced their source humans'
survey and behavioral responses 85% as accurately as the humans
reproduced themselves two weeks later.
- Simile — post-trains on interviews, transaction data, and registered
RCTs from the Open Science Framework specifically to recover human bias,
inconsistency, and causal texture, and reports early scaling laws for
simulation quality. The hurdle: frontier models are trained toward being
agent models, which makes them bad simulations of real people.
- SimGym (Shopify) — simulating shopper trajectories; the follow-up
("2,000 robots walk into a shop", Feb 2026) validated real serving
optimizations — partitioning, speculative decoding — from simulated
shoppers alone.
- Tencent's billion-persona approach — the crude end of the spectrum.
Mid-2026 was the stage's money quarter:
- Simile's valuation went vertical — $100M Series A (Feb 2026) to a $2B
valuation by July 2026 (Greenoaks-led Series B, ~20x in five months,
revenue up 5x). Synthetic subjects became an enterprise budget line.
- The competitor split — Simile requires real interviews to calibrate
each twin; rivals (e.g. Minds) generate personas from descriptions alone,
trading fidelity for instant availability on any audience.
- SimAB (arXiv, March 2026) — reframes A/B testing as fast,
privacy-preserving simulation with persona-conditioned agents.
- Delegation, not just measurement — Habermolt sends AI representatives
into public deliberation standing in for humans; Moltbook is a
Reddit-style platform populated entirely by autonomous agents with humans
only observing. The subject is simulated even when nobody asked a
question.
The focus group, the user study, the A/B panel — and now the town hall —
are becoming inference workloads.