At a fundamental level, world models do exactly what the name suggests: They build a mathematical model of the world that can then be used to make predictions about how it will change in response to certain actions or changing conditions. The "world" in this context doesn't necessarily mean the entire physical reality. Instead, it refers to the environment the model operates within, which could be anything from a warehouse to a video game.
The hope is that by developing a richer understanding of complex environments, world models could allow AI to finally break out of the chat interface, with potentially game-changing applications in areas like robotics, autonomous driving and scientific discovery.
While world models are the latest buzzword in Silicon Valley, the concept has deep roots. It first came to prominence in the 1950s, Manling Li, an assistant professor of computer science at Northwestern University, told Live Science. It arose when cognitive scientists attempted to describe the mental models people used to simulate their environments in their heads.
However, the term "world models" today refers primarily to neural networks that learn models of their environment by training on data. The modern incarnation of the idea can be traced to a 2018 study titled "World Models," by scientist David Ha and deep learning pioneer Jürgen Schmidhuber. Early models from Google, like PlaNet and Dreamer, were among the first to solve tasks by first making predictions about the outcome of different actions.
Making those predictions, Manling Li said, consists of two key tasks: state estimation and state transition. State estimation refers to the ability to perceive the current state of the environment and encode it into a format that the model can compute, while state transition means the ability to predict how a particular action will cause the environment to evolve.
The data that powers new realities
World models are trained primarily on video data, although they can also be trained on 3D data captured by light detection and ranging (lidar) or other depth sensors, audio data and even text that explains the relationships between elements in the environment. Crucially, Yunzhu Li said, this has to be paired with action data — things like robot joint angles, movement readings from an inertial sensor, or event text labels describing what action was taken.
Robots, including humanoids, will be increasingly reliant on strong world models in order to interact with the physical realm. (Image credit: China News Service via Getty Images)This data can be processed in different ways, Yunzhu Li said. One of the most popular approaches is to operate directly on raw pixel data, which represents the state of the world as a series of images and predicts how actions will change them. Another is to use the data to learn 3D geometric representations of the world that more explicitly encode spatial and physical relationships among objects in a scene.
In a pixel-based model, these abstract representations are reconstructed into pixels to make predictions about what will happen next. But it's also possible to do those simulations within the latent space by directly predicting the embedding of the environment's next state. This approach has been popularized by computer scientist Yann LeCun, Meta's former AI head and one of the "godfathers of deep learning," with his Joint-Embedding Predictive Architecture. He has raised more than $1 billion for a startup called AMI Labs, which plans to use the approach to build world models.
"As humans, when we're imagining the evolutions of the environment, we don't have to imagine the exact value of every pixel," he said. "That is why it makes a lot of sense to think about predicting over the latent space. It is easier to make sure you are only learning things that are task relevant and ignoring the things that are irrelevant to the task."
Yann LeCun, the executive chairman of AMI Labs, has raised more than $1 billion to build world models. (Image credit: Bloomberg via Getty Images)But questions about how best to represent data in a world model are secondary to the bigger issue of where to get that data in the first place, Manling Li said. LLM makers could simply scrape all the text from the internet for the initial foundational models, but high-quality, action-labeled video often has to be painstakingly curated. What's more, that data can be very sparse, she added, because only a small number of pixels in an image may change in response to an action.
This is leading to considerable debate about the best architectures for world models. Almost every LLM today is based on the transformer architecture, which excels at rapidly ingesting huge amounts of data. But these models are tuned to dense language data where every word carries some meaning, and they are less suitable for sparse video data, Manling Li said. As a result, people are experimenting with a wide variety of model architectures and the field has yet to converge on a tried-and-true recipe.
The evolution of world models
"It is only conditioned on some initial language prompt and then predicts the entire video," he said. "So it cannot predict the counterfactual futures — for example, what would have happened if you applied a different action?"
Driverless cars are one kind of AI-powered device that stand to gain from more sophisticated world models. (Image credit: Heather Diehl via Getty Images)
"In order to learn the most effective word models, it's highly likely we will also need the world model to make interactions with the environment and learn from those online interactions," Yunzhu Li said. That remains a stretch goal, however, as current neural network technology is incapable of this kind of continual learning. In addition, allowing a half-finished model to interact with the real world raises significant safety concerns, Yunzhu Li added.
Related storiesEven a more modest world model could prove invaluable for a host of applications, though. Some of the most obvious include helping robots and autonomous vehicles navigate and plan how to complete tasks. But they could also act as a general-purpose simulator for a variety of applications, depending on the data they are trained on, Manling Li said. Such simulators could include more advanced physics engines for video games, digital twins of patients that could guide medical treatment, or even new ways to model physical phenomena like the climate.
"People are working very hard and hope that with enough compute, with enough data, and with good enough algorithms, we will have this one unified world model that works across the board for many different applications," he said.
Hence then, the article about world models are the future of ai but how do they work was published today ( ) and is available on Live Science ( Middle East ) The editorial team at PressBee has edited and verified it, and it may have been modified, fully republished, or quoted. You can read and follow the updates of this news or article from its original source.
Read More Details
Finally We wish PressBee provided you with enough information of ( 'World models' are the future of AI, but how do they work? )
Also on site :
- Best-Selling Author and Fan-Favorite Comic-Con 2026 Guest Dominates Amazon Charts With Beloved Series
- “Chicharito” Hernández tiene nuevo club: lo que debes saber del Atlético Dallas y el USL Championship
- Kaitlan Collins addresses Trump’s WHCD insults and tells viewers to focus on the president’s ‘non-answers’