The Robots Cometh

Achieving this milestone will require overcoming two self-reinforcing problems: the scarcity of 3D data and the diversity of deployment environments. Unlike LLMs like ChatGPT, which can parse the virtually boundless internet, robots need manipulation data collected in the real physical world, which is costly and difficult to scale. It must typically be collected manually by repetitively training machines how to perform certain tasks, such as folding shirts, stocking refrigerators, or shaking a mai tai. Plus, explains Jiang Zheyuan, CEO of Beijing-based humanoid firm Noetix, “even if a well-performing manipulation strategy is trained for a specific task, the success rate often drops sharply when the environment changes slightly—for example, object displacement, different lighting, or different tabletop materials.”
Thanks to firms like Unitree, the hardware for humanoids is advancing quickly. The real battleground is over the robot brains, or vision-language-action (VLA) models, which are designed to turn human inputs into embodied AI performing intuitively and accurately. The VLA space alone will be worth $40.50 billion by 2035, predicts Kaiso Research, a market intelligence firm, and many of tech’s biggest players are aggressively pursuing their own models, including Nvidia’s GR00T N1.7, Google DeepMind’s RT-2-X, and Alibaba’s Qwen Robot Suite.