- General Intuition, which trains AI on video game footage to control robots, is in talks to raise at a 6 billion dollar valuation, up from 2.3 billion just eight weeks earlier.
- The round is part of more than 3 billion dollars that has moved into world models in 2026, the systems that learn how physical space and time behave rather than how language reads.
- The bet is that the next leap in AI comes not from bigger chatbots but from models that can act in the real world, a problem today's large language models cannot solve.
General Intuition's valuation nearly tripled in eight weeks
General Intuition is not a household name, and eight weeks ago it was worth 2.3 billion dollars. It is now in talks to raise at roughly 6 billion, with Valor Equity Partners, Point72 Ventures, and Seven Seven Six joining existing backers Khosla Ventures and General Catalyst, according to PitchBook data on the company. The startup spun out of Medal, a platform that stores short clips of people playing video games, in late 2025. Its founder, Pim de Witte, is turning that archive into training fuel for a foundation model that learns how objects move through space and time, then uses that understanding to control robots.
The speed of the markup is the tell. Investors are not pricing revenue, because there is almost none. They are pricing a thesis: that the technique General Intuition is chasing is the missing piece between AI that talks and AI that acts.
What a world model is, and why language models cannot do it
A large language model predicts the next word. It has read most of the internet, but it has never touched a cup, judged a gap in traffic, or felt an object slip. A world model predicts the next state of an environment instead of the next token. Given a scene and an action, it forecasts what happens next in physical terms, which lets a system plan, and act, rather than merely describe. That is the capability robots have always lacked and chatbots were never built to provide.
Three efforts define the field, and they are converging from different directions. General Intuition mines gameplay, a near infinite, physics consistent record of cause and effect. Fei-Fei Li's World Labs raised about 1 billion dollars to build what it calls spatial intelligence, models that generate and reason about full 3D worlds, shipped in a product called Marble. Nvidia took the infrastructure route, releasing its open Cosmos world foundation models for physical AI to any robotics team that wants them.
| World-model effort | Founder or backer | Approach and signal |
|---|---|---|
| General Intuition | Pim de Witte, ex-Medal | Learns physical dynamics from video game footage to control robots. About 6 billion dollar valuation in talks, from 2.3 billion in June 2026 |
| World Labs | Fei-Fei Li | Spatial intelligence models that generate and reason about 3D worlds. Raised about 1 billion dollars, product Marble |
| Nvidia Cosmos | Nvidia | Open world foundation models for physical AI and robotics. Cosmos 3 released in June 2026 |
Sources: PitchBook and company disclosures on world-model and physical-AI efforts, 2026.
The training data advantage hiding inside video games
The non-obvious part of General Intuition's pitch is not the model, it is the fuel. Robotics has always been starved of data, because collecting real world motion is slow, expensive, and hard to label. Video games solve that quietly. Every clip is a labeled simulation of physics, intent, and consequence, produced for free by millions of players, and General Intuition sits on Medal's library of them. In a field where the modeling ideas travel fast, a proprietary supply of physics consistent, action labeled footage is the kind of moat that is difficult to copy.
Language models learned to read the internet. World models are learning to move through it. The company that owns the best simulation of reality may end up owning the robots that operate in it.
What shifts if the frontier moves from language to physics
For robotics, a working world model is the difference between a machine that repeats one trained task and one that generalizes to a kitchen, a warehouse, or a road it has never seen. That is the prize, and it explains why capital is arriving at chatbot scale valuations before a single robot ships at volume. It also explains the risk. The gap between a model that predicts a simulated world and a robot that behaves safely in a real one, the sim to real gap, has humbled the field for a decade, and none of these companies has closed it yet.
The strategic read is that the AI map is being redrawn while attention stays fixed on chatbots. The incumbents best positioned are the ones that already own simulation, sensors, or robots: Nvidia through Cosmos, and the robotics and autonomy programs at Google, Tesla, and a wave of humanoid startups. The vulnerable ones are pure language labs whose advantage does not obviously transfer to a domain where being fluent counts for nothing and being right about gravity counts for everything.
For three years the race in AI was a race to read and write. General Intuition's valuation, and the capital stacking behind spatial models, is a bet that the next race is to move. If it pays off, the most valuable AI companies of the coming decade will not be the ones that answer questions. They will be the ones that understand rooms, roads, and hands well enough to let a machine act inside them.
Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.