More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Physical AI struggles because gathering robot data costs real money and time, unlike scraping text. Manual teleoperation scales linearly with labor—Goldberg’s estimate of 100,000 human-years to hit frontier performance makes that clear. Vendors push raw hours as their key metric, but hours spent often correlate poorly with model gains. Deploying robots into paid production seems cheap on paper, but those environments are low-variance. They churn out repetitive telemetry that adds little new signal.
The authors break robot data into three classes: observational video with no action labels, teleoperation demonstrations rich in state-action pairs, and deployment telemetry from live systems. Observational data gives breadth cheaply but needs action supervision added later. Teleoperation is precise but pricey. Deployment data arrives “free” with revenue but is often narrow and low-entropy. They borrow scaling-law insights from language models: loss drops as a power law with data size, but diversity and novelty drive that exponent. Repeating the same scenario more than four times yields almost no extra utility; beyond sixteen passes, it can even hurt.
To squeeze value per dollar, the paper argues, you need to price data by novelty—not volume. Drawing on joint scaling formulas (Kaplan 2020; Hoffmann 2022) and data-mixing results (Ye et al. 2024), they show that distinct, cross-domain examples push down the irreducible error floor faster. Neo-integrators like Standard Bots bet that live deployment will fuel a learning flywheel. Critics like Kyle Vedder warn those early niches won’t supply enough fresh scenarios. The real capital play lies in balancing spend across observation, teleop, and deployment—tilting toward sources that expand the model’s intrinsic dimension instead of drowning it in redundant logs.
Questions about this article
No questions yet.