The End of the Demo Robot Era
If you've been to a robotics expo lately, you might have noticed something odd. The flashy robots that used to steal the show—the ones that wave, dance, or do backflips—are getting less attention. Now you see robots sorting parcels, checking pipes, folding shirts. It's less about spectacle, more about getting stuff done.
At WAIC, over 200 embodied AI companies were showing off their hardware. The message? Robots are moving from "look at me" to "put me to work." But what makes a robot that can perform a trick different from one that can actually help in a gym or a warehouse? It's not the arms or the wheels. It's the brain.
Why Physical Work Needs a New Kind of AI
For a long time, AI researchers used a simple trick: take a model trained on internet data and tweak it for a new task. That worked for language and images. But for robots, that shortcut doesn't cut it.
Yujun Shen, chief scientist at Ant Lingbo, is blunt about this. He calls it a "path dependency"—a convenient detour, not the real deal. His team is building models from scratch, designed specifically for physical interaction. No borrowed brain. No repurposed tricks.
Why? Because the physical world is a mess. It's noisy, unpredictable, and full of rules that no amount of YouTube videos can teach.
Data: The Uncharted Frontier
Here's a dirty secret: there's no universal standard for robot training data. One company uses a single head-mounted camera. Another swears by a five-finger gripper with sensors. Some have humans puppeteer the robot to show it what to do.
So the data is all over the place. Do you need just visual input? Head position? Hand skeleton data? Full hand pose with touch? Nobody agrees. And because the model architectures aren't settled either, it's a chicken-and-egg problem.
Data collection is stuck in the Wild West. Every team is guessing which dimensions matter most, and no one wants to compromise on quality because they don't know what they'll need later. It's inefficient, expensive, and honestly, a bit chaotic.
Simulation vs. Real Data: A Trade-off
Some researchers think simulation is the answer. Train a robot in a virtual world, generate endless scenarios without breaking a sweat—or a robot. That works for narrow tasks like autonomous driving, where the environment is predictable.
But for general-purpose robots—the kind that might one day help you set up a badminton net or fetch a water bottle—simulation has limits. "Simulation can't capture how a human instinctively opens a bottle," Shen notes. "It can't model the randomness of a ball bouncing off a table."
That's why Shen's team leans heavily on real data, even though it's harder to collect. They focus on two types: teleoperated data from actual robots, which teaches the robot about its own body, and first-person data from humans, which captures authentic human behavior. Internet data? Fine as a base, but it can't fill the gaps.
VLA vs. VA: Two Roads, Same Goal?
In robotics research, you'll see acronyms like VLA (Vision-Language-Action) and VA (Vision-Action). VLA models are great at understanding commands and switching tasks. VA models, which come from video generation, handle randomness better.
Which will win? Shen says neither—at least not yet. "Both have clear weaknesses," he explains. VLA is fast and robust to messy data, but it struggles with unpredictable environments. VA is good at probabilistic reasoning, but it's picky about data quality and slower to adapt.
Instead of betting on one, Lingbo is running both. They released six models at once, each tackling a piece—vision, depth, video prediction, action. It's like building a robot brain from Lego blocks, snapping them together as each piece proves itself.
Safety Can't Be an Afterthought
Ask any robotics engineer about the biggest barrier, and they'll mention safety. But current approaches are mostly "fences"—rules that tell the robot what not to do. Don't get too close to the edge. Don't knock over the glass. That works in controlled settings, but it falls apart in the real world.
Picture a robot pouring water. Seems harmless. But what if there's an electrical outlet nearby? The robot doesn't know that pouring water near a socket is dangerous because it's never seen that exact scenario. You can't pre-program every possible hazard.
Shen argues for "native safety"—building safety awareness into the model from the start, not bolting it on later. It's a tall order, but he thinks it's essential. "We can't wait until robots are already in homes to figure this out," he says. "We need to start now."
The Road to a Robot GPT Moment
Everyone's waiting for robotics to have its ChatGPT moment—that breakthrough where the tech suddenly feels useful. But Shen thinks the path there might be less glamorous than we imagine.
For language models, the spark came from the chat interface. For robots, he believes it will come from data. Imagine if everyday people could easily contribute robot training data—just by doing a task at home or at work, like how Tesla cars collect data as owners drive. That could be the turning point.
"When a regular person can spend an hour helping a robot learn, and the robot gets better because of it, that's when robots start to feel real," Shen says.
Until then, the industry will keep grinding away on the unglamorous work: standardizing data, refining models, and teaching robots not just to move, but to understand the physical world. It's not as flashy as a robot doing backflips, but it's the foundation everything else will stand on.
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!