World models take an observation and an action and return a predicted state, not the next word — the layer today's chatbots lack, and the foundation for 'physical AI': robots and other embodied systems that have to act in the real world rather than
Yann LeCun and Fei-Fei Li are walking away from text-only AI and putting their chips on a different class of model: "world models."
A world model takes an observation of an environment plus a proposed action, then returns an embedding predicting what the world looks like after the action. A home robot asked to light a stove, for example, would get back a predicted state of the kitchen, not the next word in a sentence. The unit of work is a state, not a token, which is what makes the approach a candidate for "physical AI," the loose industry term for embodied systems that have to act in warehouses, labs, homes, and roads rather than only chat.
A Forbes Technology Council piece by Yifeng Yin, CEO of TEA Intelligent System, surfaces the schism to a business-tech audience and lays out the textbook distinction. This excerpt cites LeCun and Li second-hand; their primary papers were not re-fetched for this report, so the public case for world models here rests on contributed framing rather than a primary survey.
The category now has a name. Who can train one well enough to land a real robot in a real kitchen is the open question.