For two years, the entire field has optimized for one thing: making models think longer.
More reasoning tokens, more test-time compute, longer chains of thought, agents arguing with other agents.
It worked. OpenAI’s o1 line proved that inference-time computation is its own scaling axis, independent of model size. Give a model more time to work through a hard problem and it gets measurably better at it.
But that’s not where the next jump comes from. The next jump comes from teaching AI to simulate the world instead of just reasoning about it—and that distinction is bigger than it sounds.
Ask a reasoning model whether you should cut your price 20%, and you’ll get a genuinely useful argument: lower prices should lift conversion, margins will compress, competitors might retaliate, so maybe 10% is the safer move. That’s real reasoning—the model manipulates concepts and relationships until it lands on a conclusion.
But it’s still one argument, generated once.
A simulator does something structurally different. It builds a state—your customers, competitors, cost structure, acquisition channels, and market conditions—intervenes on it (price = −20%), and rolls the environment forward.
Run it once and revenue rises 4% while margin drops 18% and a competitor reacts in six weeks.
Run it again and it’s a price war.
Run it 100,000 times with different assumptions and you stop getting an opinion and start getting a distribution: here’s what happens, how often, and under which conditions each strategy actually wins.
This isn’t a new idea. It’s an old one finally converging.
The pattern has been building in AI for a decade, just not inside language models.
AlphaGo paired a policy network with a value network so it could search possible futures instead of only pattern-matching to prior games. AlphaZero pushed further by generating its own training data through self-play rather than imitating humans.
Then MuZero removed the last dependency: it didn’t need to be told the rules of Go, chess, or Atari. It learned an internal model of the environment and planned by searching inside that learned model—proof that a simulator doesn’t need to reconstruct every detail of reality, only the details that matter for the decision at hand.
LLMs arrived on a different track entirely, built on language prediction rather than world dynamics.
Reasoning models extended that track by spending more computation on intermediate steps before answering.
But notice what’s actually being searched in a reasoning model: possible paths through thought space. A simulator searches through future state space. That’s a different primitive, and it’s why the two tracks are now colliding rather than one simply replacing the other.
Why reasoning alone hits a wall
A reasoning model can produce a convincing explanation and still be wrong about what actually happens. A plausible chain of thought is not the same as a reliable model of reality.
Simulation doesn’t magically solve this—a bad simulator can hallucinate the future too. But it changes what we can test. Instead of asking only “Was the answer correct?”, we can ask “Did the predicted world evolve like the real one?”
That’s a much stronger signal.
Why this matters
Reality gives you one attempt.
You launch one product, hire one person, make one investment.
Simulation could let AI test thousands of possible futures before taking the real action.
Instead of asking “Is this startup a good investment?”, you could simulate the company across different markets, competitors, hiring decisions, and funding environments.
Instead of asking whether a feature will work, simulate users interacting with it first.
Coding agents already do a primitive version of this: they don’t just reason about code, they execute it, test it, observe failures, and try again.
The biggest problem
A simulator is only useful if its world resembles reality.
Small errors compound. Humans and competitors react unpredictably. Models can exploit flaws in their own simulated environments.
So the key question for world models isn’t “How realistic does the simulation look?”
It’s “How well does success inside the simulation predict success in reality?”
The future of AI probably isn’t reasoning or simulation.
It’s reasoning deciding which futures to simulate, simulation showing what might happen, and agents choosing the future most likely to survive contact with reality.