The Problem with Secondhand Brains
Walk into any tech conference these days and you'll see robots. Lots of them. But look closer and you'll notice something odd: most of them are just performing. They wave, they dance, they maybe pick up a cube. It's like watching a highlight reel of someone who learned martial arts by watching movies—all flash, no substance.
That's the gap between a robot that can look impressive and one that can actually do real work. And it's exactly the gap that Ant Group's LingBot is trying to close with what they call a "full-stack brain." Their bet is that the missing piece isn't better hardware or more sensors—it's a brain trained from scratch for the physical world, not one borrowed from the internet.
Why Video Models Don't Cut It
Here's the thing about AI: most of the progress we've seen in the last decade came from feeding models internet data. Text, images, video—all scraped from the web. That works great for generating content, but it's a terrible foundation for a robot that needs to grab a cup without crushing it.
LingBot's chief scientist, Shen Yujun, puts it bluntly: trying to adapt a digital-world video model for robotics is a "shortcut," not a solution. It's like a fighter who only trained on shadowboxing and then steps into the ring. Sure, they know the moves, but they don't have the instincts.
So what does a robot actually need? It needs to understand distance, touch, temperature, and cause-and-effect in real time. It needs to know that if it pushes a glass, the glass will fall and break. It needs to anticipate, not just react.
Enter the "Embodied Native" Model
LingBot's answer is what they call an "embodied native" model. That's a fancy way of saying the model is designed from the ground up for physical interaction, not retrofitted from a language model or a video generator.
At this year's WAIC conference in Shanghai, LingBot unveiled six new models as part of their full-stack brain 2.0. These cover everything from spatial awareness to dexterous manipulation to environmental feedback. The centerpiece is LingBot-VA 2.0, which they claim is the industry's first embodied native world-action model.
The idea is simple: don't try to cram all the complexity into one massive end-to-end model. Instead, break the problem down into manageable pieces, verify each one works, and then slowly stitch them together. It's like building a fighter's skill set—you don't start with a full MMA game. You work on your jab, your takedown defense, your cardio. Then you put it all together.
Data: The New Frontier
But here's the rub: all these models need data, and physical-world data is a mess. Unlike the internet, which has decades of text and images, there's no massive archive of robot manipulation data. Every company is collecting its own, with no standard format.
What data should you collect? Just visual? Or also depth? Do you need hand poses to the millimeter, or is centimeter-level good enough? These aren't trivial questions. The answers determine what you can train, and they're still up in the air.
Shen admits the industry is stuck in a chicken-and-egg loop: data standards depend on model architectures, and model architectures depend on data. But he's optimistic. "In the last few months, the model routes are slowly converging," he says. Once that happens, data will follow, and then the flywheel starts spinning.
VLA vs. VA: Two Roads to the Same Destination
If you follow robotics AI, you've probably heard of VLA (Vision-Language-Action) models and VA (Vision-Action) models, also known as world action models. They're two competing approaches, and people love to debate which one will win.
LingBot is doing both. Why? Because they think neither is the final answer. VLA is great at understanding language and switching tasks on the fly. VA is better at handling randomness and using data efficiently. But each has clear weaknesses. "If we simply go along these two routes, they might still have a long way to go," Shen says. "And the end might not be the end."
It's like comparing a counter-striker to a pressure fighter. Both can win, but a truly great fighter needs to blend both styles.
The Hard Part: Training from Scratch
Here's the part that doesn't get enough attention: training a model from scratch is hard. Really hard. Most teams just take an existing open-source model and fine-tune it. That's like buying a used car and changing the tires. It works, but you're not building a new car.
LingBot is trying to build the car. They've had to figure out how to train a vision model with spatial awareness, how to make a Mixture-of-Experts model actually use all its experts (turns out, most of them go idle), and how to enforce causality—the simple rule that the future can't affect the past.
For a robot, causality is everything. A video model can cheat by looking at the whole clip before generating a frame. A robot can't. It only has the present moment and the past. So LingBot had to design training methods that force the model to respect that constraint.
Safety: More Than a Cage
And then there's safety. Most robot safety today is what Shen calls "fence-based." You put a virtual cage around the robot and tell it to stop if it gets too close to a person or an obstacle. But that's like telling a fighter to avoid getting hit by never engaging. It doesn't work in the real world.
Instead, Shen argues, robots need "native safety"—a deep understanding of what's dangerous. "Safety is a higher-order intelligence," he says. It needs to be baked in during pre-training, not bolted on later.
It's a bold vision, and it might be years from reality. But if robots are ever going to step into our homes and do real work, they'll need more than just skills. They'll need judgment.
The ChatGPT Moment for Robots
Everyone's waiting for the "ChatGPT moment" for robotics—that instant when the technology suddenly feels useful to regular people. For ChatGPT, that moment came when you could just chat with it. For robots, Shen thinks it'll happen when ordinary people can easily contribute to robot data collection.
"If one day, robot data collection becomes so convenient that each person, while going to work, can spend an hour helping robots generate training data—and the robots get better because of it—that's when robots truly enter people's lives," he says.
Until then, we're stuck with robots that can dance but can't set a table. The race is on to build the brain that changes that. And it's not going to come from a video model.
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!