More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
A new wave of AI is moving off screens and into the physical world, driven by three fast-growing areas: robot learning, automated scientific discovery, and next-gen human-machine interfaces. In robot learning, models like Physical Intelligence’s π₀, Google DeepMind’s Gemini Robotics, and NVIDIA’s GR00T N1 extend image-text backbones into action decoders, while systems such as NVIDIA’s DreamZero use video-trained transformers to predict object dynamics and generate motor commands. A third path, represented by Generalist’s GEN-1, skips internet data altogether, training on half a million hours of real-world interaction captured from wearable sensors. Each approach aims to compress the physics of everyday manipulation into reusable models.
Behind these efforts lie shared technical building blocks. First is learned representations of physical dynamics—models that predict how objects move, deform, and respond to force. VLAs borrow semantic features from large vision-language networks, WAMs use video diffusion to absorb motion priors, and native embodied models ingest direct human-object contact data. A missing piece in all three is true 3D spatial reasoning. Companies like World Labs step in with spatial intelligence systems that reconstruct full scene geometry, lighting and layout, filling the gap left by 2D-only inputs.
The next layer is architectures for embodied action—software that turns high-level goals into stable, continuous motor outputs over long time horizons. On top of that sits closed-loop orchestration: synchronizing perception, planning, and control in real time. Finally, simulation and synthetic-data pipelines feed these models with scalable, diverse training experiences. When you combine robot arms folding towels, self-driving labs predicting chemical reactions, and brain-computer interfaces decoding motor cortex signals, you get a feedback loop: more data sharpens dynamics models, which improve action planning, which in turn accelerates real-world deployment and new data collection. That loop could be the key to unlocking AI’s next big leap—directly manipulating atoms, materials and minds.
Questions about this article
No questions yet.