Where Simulation Breaks (2026)
Every stage of this grid has a matching 2026 correction paper. The
skeptical literature is itself a trend worth tracking: it maps exactly the
residue Stage 8 predicts — the part of "human" that doesn't compress.
The variance failure
Independent evaluations of synthetic respondents converge on one finding:
- Simulated panels track the average response reasonably well, then fail
on variance, price sensitivity, and distribution tails — precisely
where high-stakes decisions live.
- Fully synthetic focus groups break on lived experience and emotional
nuance: the inputs that cannot be rephrased from text.
A simulated population that gets the mean right and the tails wrong is not
a cheaper focus group; it is a different instrument that happens to share a
mean.
Model collapse moves downstream
The collapse argument — models trained on model output regress toward the
mean of their training distribution — now applies to self-feeding
simulation loops, not just pretraining corpora:
- A judge trained on judgments of model outputs drifts toward what models
find judgeable.
- Environments synthesized from model-mined "work patterns" encode the
model's idea of work.
- Personas distilled from personas lose the very inconsistency Simile
spends interview data to restore.
The loop closing on itself is also the loop narrowing on itself.
The "AI slop" correction
Consumer-facing synthetic media hit an uncanny-valley backlash in 2026 —
campaigns measurably losing sentiment when audiences perceived content as
synthetic. The market response, "human-in-the-loop," is an admission that
full substitution failed at the last mile of being received as human.
What practitioners take from this
- Match the instrument to the question — synthetic panels for
directional average effects; real humans for variance, price, tails, and
anything with legal or health stakes.
- Calibrate against ground truth continuously, the way Simile anchors
twins to interviews and RCTs — simulation quality is a measured quantity,
not a property.
- Stress-test the simulators, the way synthesized verifiers are
stress-tested before training: oracle, no-op, and unsolved probes — for
environments and for populations.
This is the layer the Simulacra playground models: verifiers earn trust by
failing the right attempts, and simulated populations should be held to the
same standard before anyone acts on their output.