Robots Are Learning to Grip by Watching Human Hands

Instead of teaching robots with robot data, one new model trained on 100,000 hours of humans using a handheld camera-and-gripper rig.

Robots Are Learning to Grip by Watching Human Hands

Robotics has a data problem that language models never had. Text is everywhere; footage of a robot arm competently loading a dishwasher is not. Every hour of training data for a physical task has historically meant an hour of an actual robot, in an actual lab, slowly and expensively failing at something a toddler would manage on the second try.

The workaround gaining traction this year is disarmingly simple: stop waiting for the robot. Xiaomi-Robotics-1, the company’s new foundation model, was trained on more than 100,000 hours of motion data collected by people using camera-equipped handheld grippers — essentially a tool that looks like a set of salad tongs with a camera bolted on, letting a human demonstrate a task while recording exactly what a gripper “sees” and how it moves.

It’s a clever sidestep of the classic chicken-and-egg problem in embodied AI. You don’t need a fleet of humanoids to generate humanoid-relevant training data; you need a cheap sensor and a lot of people doing ordinary tasks. As one write-up of the release put it, adding data improved performance far more than increasing model size did. The hope is that motion and grip patterns captured this way transfer reasonably well to an actual robotic hand, the same way pretraining on internet text transfers to tasks nobody wrote a dataset for.

Whether this scales to genuinely dexterous, general-purpose manipulation is still an open question — human hands and robot grippers are not the same instrument, and the sim-to-real gap has humbled plenty of promising approaches before. But as a way of getting the first order of magnitude of data cheaply, it’s a pragmatic bet, and one several other robotics teams are likely to copy before the year is out.

Sources