Every model in this series so far lives on a screen. This one is starting to walk around. This is the fifth and final article in LIWARSE’s series on the AI landscape by type of use — and the one where LIWARSE’s Artificial Intelligence and Robotics pillars meet directly.
The dominant architecture behind today’s robots is the vision-language-action (VLA) model — a system that takes in camera images and a plain-language instruction and outputs motor commands directly, the same way a chatbot takes in a prompt and outputs text. Over 50,000 humanoid robots are estimated to be operating commercially by the end of 2026, up from roughly 16,000 a year earlier — and nearly all of them run on some version of this design.
The Model Builders: Google DeepMind, NVIDIA, and Physical Intelligence
Google DeepMind’s Gemini Robotics line, built on the Gemini architecture, adds 3D spatial perception and the ability to generate robot control code on the fly; an on-device release made the model lightweight enough to run locally on the robot itself rather than depend on a network connection — a meaningful safety property for any robot operating where connectivity cannot be guaranteed, including hospitals and disaster zones. NVIDIA’s Isaac GR00T family pairs a vision-language reasoning backbone with a diffusion-based motor-control module, released as an open, customizable foundation model that any humanoid maker can fine-tune — the clearest case of an “open-weight” release in the physical-AI category. Physical Intelligence’s π0 and π0.5 models, meanwhile, have demonstrated some of the most general household manipulation shown publicly: tidying kitchens, bathrooms, and bedrooms the model has never seen before.
The Hardware Makers Training Their Own Brains
A second group of companies builds the robot and the model together in-house. Figure AI’s Helix model has shown two humanoid robots coordinating on a shared packing task using natural language, and its newer Figure 03 platform integrates deeply with OpenAI. Tesla’s Optimus program runs an internal VLA derived from the same architecture family as Tesla’s self-driving software, aiming — though not yet delivering — a consumer price point near $20,000–$30,000 by 2028. Boston Dynamics’ Atlas, Agility Robotics’ Digit, Apptronik’s Apollo, and Unitree’s G1 (the most affordable at roughly $16,000–$23,000) round out a field moving from single-task pilots to multi-hour factory shifts on real production lines, largely in automotive sub-assembly.
Beyond humanoids, the same VLA recipe is being applied to warehouse manipulation (Amazon Robotics’ Covariant-derived systems), surgical robotics research (Intuitive’s da Vinci 5 platform), and self-driving (Tesla FSD, Wayve) — evidence that this is a general-purpose architecture for embodied action, not a humanoid-specific trick.
Risks and Benefits Through the LIWARSE Lens
Benefits
- On-device models like Gemini Robotics’ local release reduce dependence on network connectivity, which matters directly for robots assisting in hospitals, disaster response, or any setting where a dropped connection could otherwise leave a machine mid-task with no oversight.
- Open foundation models like NVIDIA’s Isaac GR00T lower the barrier for smaller robotics labs and researchers to build safety-relevant applications — assistive and care robotics among them — without training a foundation model from scratch.
- Generalist manipulation models (π0.5) that can operate in homes they have never seen point toward practical eldercare and disability-assistance robotics sooner than task-specific programming ever could.
Risks
- An open, downloadable robot foundation model (Isaac GR00T) is a materially different hazard from an open chatbot: the “output” is physical force in a shared human space, not text on a screen, which is exactly why LIWARSE’s four-layer containment model treats agentic and embodied AI as a distinct, higher-stakes category.
- Robotics remains data-starved by orders of magnitude compared with language models — the largest known robot datasets have roughly a billion timesteps against an LLM’s tens of trillions of tokens — which means these models generalize far less reliably than their language-model cousins, a gap easy to underestimate given how fluently VLA demos read.
- Full-stack humanoid makers training proprietary in-house models (Tesla, Figure) face no equivalent of the safety-tuning transparency debate that at least exists, however imperfectly, for open-weight language models.
The LIWARSE Assessment
This category is where LIWARSE’s four-layer containment model — compartmentalization, sandboxing, specialist oversight, and tiered access — was written for. A hallucination from a chatbot is an embarrassment to correct. A hallucinated action from a humanoid robot moving boxes near a person is a physical event that cannot be undone by regenerating the response. As commercial deployment accelerates from roughly 16,000 to over 50,000 units in a single year, the industry’s safety infrastructure — verification, fail-safes, human-on-the-loop authority — needs to accelerate at least as fast as the hardware does, not follow it.
Under the 3 Absolute Laws, no robot foundation model — however capable — is exempt from human oversight and shutdown authority simply because it moves too fast for a person to double-check in the moment. That is precisely the moment oversight matters most.
— The LIWARSE Movement | liwarse.org
Safety of Life · Advancement of Life · Together.