On this page
The humanoid brain
"Whether we are based on carbon or on silicon makes no fundamental difference; we should each be treated with appropriate respect." - Arthur C....
"Whether we are based on carbon or on silicon makes no fundamental difference; we should each be treated with appropriate respect." - Arthur C. Clarke
A humanoid robot reaches for a coffee cup on a cluttered kitchen counter. To a human, this takes less than a second and requires no conscious thought. For the robot, it triggers a cascade of computation that reveals everything difficult about building minds for machines.
First, cameras capture the scene. But pixels are not objects. The robot's vision system must segment the image, identifying the cup among dozens of other items, estimating its position in three-dimensional space, recognizing that the handle faces left, that the liquid inside might slosh if grasped too quickly. This alone requires processing millions of data points against learned models of what cups look like, how they behave, where handles tend to be.
Then comes planning. The arm must chart a path from its current position to the cup, avoiding the salt shaker, the stack of mail, the edge of the counter. The path must be smooth enough to look natural, fast enough to be useful, and safe enough that a collision won't send ceramic shards across the floor. The hand must rotate to align with the handle, fingers must close with precisely enough force to grip but not crush.
Finally, execution. Motors fire. Joints rotate. Sensors monitor progress, feeding back data that adjusts the plan in real time. The cup is heavier than expected. The surface is slick. The robot compensates, recalculates, adapts.
All of this happens in the time it takes you to blink. And all of this requires hardware and software working together in ways that the computing industry is only beginning to understand.
Building a brain for a smartphone means optimizing for battery life and screen responsiveness. Building a brain for a humanoid means solving a problem evolution spent millions of years on: closing the loop between perception, thought, and action in a body that moves through physical space.
The Architecture Question#
In the 1990s, PC manufacturers built their systems around the "Wintel" duo: Microsoft Windows as the operating system and Intel's x86 processors as the CPU. The architecture was settled. Companies competed on performance, not on fundamental design. A similar pattern emerged with smartphones: ARM processors, mobile operating systems, standardized sensors. The winners were those who executed best within a known framework.
Humanoid robotics has no such settled architecture. The machines entering factories and homes in 2025 represent competing bets on how to organize computation for embodied intelligence. Some companies believe the answer lies in massive centralized processors running large AI models. Others bet on distributed systems where intelligence is spread across the robot's body. Still others pursue neuromorphic chips that mimic biological neural networks rather than traditional digital logic.
What every humanoid brain must accomplish, regardless of architecture, is the continuous cycle that roboticists call the perception-action loop. Sensors feed data to processors. Processors build models of the world, plan actions, and send commands to actuators. Actuators move the body, changing the robot's relationship to the environment. New sensor data arrives. The loop repeats, dozens or hundreds of times per second, without interruption.
This loop imposes brutal constraints. Latency kills. A robot that takes 500 milliseconds to react to an obstacle will collide with it. A robot that takes 100 milliseconds to adjust its grip will drop the cup. The brain must be fast enough to keep pace with the physical world, powerful enough to run sophisticated AI models, and efficient enough to operate on battery power for hours.
No one has solved this problem completely. But several approaches are emerging, each with different tradeoffs.
The Hardware Layer#
The physical brain of a humanoid robot centers on a System-on-Chip, or SoC: an integrated circuit combining CPUs, GPUs, memory, and specialized accelerators onto a single piece of silicon. Unlike the general-purpose processors in laptops, these chips are optimized for the specific workloads of robotic intelligence: processing camera feeds, running neural networks, coordinating motor control.
NVIDIA has emerged as the early leader in this space. Its Jetson Thor SoC, built on the Blackwell architecture, delivers up to 2,070 FP4 teraflops of AI performance specifically tailored for humanoid robots. But raw compute is only part of NVIDIA's strategy. The company has built an ecosystem around its hardware: Isaac Sim for robotic simulation, Omniverse for creating virtual training environments, and Project GR00T as a foundation model for humanoid behavior. Companies including Figure AI and Boston Dynamics have adopted NVIDIA's platform, giving it significant influence over how the industry develops.
Qualcomm approaches the problem from a different angle. Known for the Snapdragon processors that power most Android smartphones, Qualcomm brings expertise in low-power, high-efficiency computing. Its Qualcomm Aware Platform optimizes real-time data processing for robotics, while its experience with edge computing, running AI models on devices rather than in the cloud, translates directly to robots that must think locally rather than waiting for server responses. Qualcomm's partnerships with HPE and Lenovo on smart edge servers suggest ambitions beyond individual robots to the infrastructure that supports them.
Intel, once dominant in computing, now competes on multiple fronts. Its Xeon server processors handle heavy AI training workloads, while its Gaudi3 accelerator targets inference, the process of running trained models to make predictions. More intriguing is Intel's work on neuromorphic computing. The Loihi 2 chip mimics the structure of biological neurons, processing information through spikes of activity rather than continuous signals. Neuromorphic systems promise dramatic improvements in energy efficiency for certain tasks, particularly sensory processing and pattern recognition. Whether this approach can scale to full humanoid intelligence remains an open question.
Arm occupies a foundational position. Rather than manufacturing chips directly, Arm licenses its energy-efficient architectures to other companies. Its Cortex CPUs and Mali GPUs appear in countless robotic systems, while its Neoverse platform targets high-performance computing for AI workloads. Arm's licensing model means its technology could end up in humanoid brains regardless of which company ultimately dominates the market.
Beyond these established players, a generation of startups is pursuing radical alternatives. Groq has developed Language Processing Units optimized for ultra-fast AI inference, achieving speeds of over 1,300 tokens per second for large language models. While designed primarily for text, Groq's approach to low-latency processing could prove valuable for robots that must reason quickly. The company raised $640 million in August 2024, valuing it at $2.8 billion.
Etched takes an even more aggressive approach. Its Sohu chip burns the transformer architecture, the foundation of modern AI, directly into silicon. By sacrificing the flexibility to run other types of models, Etched achieves extreme efficiency for transformer-based inference. For humanoids running vision-language-action models built on transformers, this specialization could dramatically reduce power consumption and cost.
Cerebras Systems has built the world's largest chip. Its Wafer-Scale Engine uses an entire silicon wafer rather than cutting it into individual processors, yielding 4 trillion transistors and 900,000 AI-optimized cores. The WSE-3 delivers performance that dwarfs traditional systems for both training and inference, though its size and power requirements limit it to data center applications rather than onboard robot use.
Lightmatter pursues photonic computing, using light instead of electrons to perform calculations. Its Passage chip integrates photonic cores with standard silicon, offering up to 10x energy efficiency for AI workloads. Light travels faster than electrical signals and generates less heat, advantages that could prove decisive as humanoid AI models grow more demanding.
SiMa.ai focuses specifically on edge AI for robotics. Its MLSoC platform delivers 50-200 teraflops for computer vision and inference while maintaining the low power consumption essential for battery-operated robots. The company raised $270 million in April 2024, signaling investor confidence in dedicated robotic AI hardware.
The diversity of approaches reflects genuine uncertainty about the right path forward. Traditional digital processors offer flexibility and mature software ecosystems. Neuromorphic chips promise biological efficiency. Photonic systems could break through power and speed limitations. Specialized accelerators sacrifice generality for performance. The winning architecture may combine elements of several approaches, or a breakthrough could render current bets obsolete.
Vision-Language-Action Models#
The software that animates humanoid robots has undergone a revolution. For decades, roboticists wrote explicit programs: if the sensor reads this value, move the motor by that amount. These hand-coded systems worked in controlled environments but failed in the messy real world, where infinite variations in lighting, object placement, and unexpected obstacles defeated every rule the programmers could anticipate.
The new paradigm is learning. Rather than programming specific behaviors, engineers train AI models on vast datasets of examples. The models learn patterns that generalize beyond their training data, enabling robots to handle situations their creators never explicitly anticipated.
The most significant development in this space is the emergence of vision-language-action models, or VLAs. These systems unify three capabilities that were previously separate: understanding what cameras see, interpreting natural language instructions, and generating physical actions. A VLA can watch a human demonstration, hear the command "pick up the red cup," and translate both inputs into the precise motor commands needed to accomplish the task.
The architecture builds on transformers, the same technology underlying ChatGPT and other large language models. But where text models predict the next word in a sequence, VLAs predict the next action in a physical trajectory. They process visual tokens representing what the robot sees, language tokens representing instructions or context, and action tokens representing joint positions or motor commands. The model learns to map from perception and language to action through exposure to millions of examples.
Physical Intelligence has developed π0 (pi-zero), a multimodal foundation model that integrates language, vision, and action into a single system. π0 enables robots to perform tasks like folding laundry or assembling furniture with human-like dexterity and reasoning. The system can accept voice commands, watch demonstrations, and generalize to novel objects and environments. Physical Intelligence raised $470 million through 2024, including a $400 million Series A that valued the company at $2.4 billion, with backing from Jeff Bezos, OpenAI, and Sequoia Capital.
Figure AI has built Helix, its own vision-language-action system for the Figure 02 and Figure 03 robots. Helix unifies perception, language understanding, and motor control, enabling complex manipulation tasks in unstructured environments like BMW's Spartanburg plant in South Carolina where Figure robots work. After initially collaborating with OpenAI, Figure developed Helix independently, giving the company full control over its AI stack.
1X Technologies created Redwood, a 160-million-parameter transformer model that combines vision, touch, and body movement data. Trained on real-world teleoperation, Redwood enables the NEO robot to chain together multiple skills via voice commands. The model learns both low-level motor skills and high-level task sequences, allowing natural language to translate directly into physical behavior.
Sanctuary AI takes a different approach with its Carbon AI control system. Rather than a single end-to-end model, Carbon integrates large behavior models for reasoning and planning with specialized modules for perception and motor control. This architecture allows Sanctuary's Phoenix robot to adapt to new tasks with minimal retraining, learning from teleoperation demonstrations in hours rather than weeks.
Tesla leverages its autonomous vehicle expertise for Optimus. The robot's AI shares neural network architectures with Tesla's Full Self-Driving system, adapted for humanoid form factors. Tesla's Dojo supercomputer, built for training FSD models, now accelerates Optimus development. Unlike most competitors, Tesla plans to keep its technology proprietary, following Apple's model of vertical integration rather than licensing to others.
Beyond these robotics-focused companies, general AI leaders are positioning for embodied intelligence. OpenAI, while primarily focused on digital AI, participated in Figure AI's $675 million Series B funding round through the OpenAI Startup Fund and continues research into embodied systems. The reasoning capabilities of GPT-5 and its successors could eventually enable robots to understand complex instructions and plan multi-step tasks in ways current VLAs cannot.
Specialized firms are building the infrastructure that VLA development requires. Skild AI focuses on lifelong learning and adaptability, enabling robots to improve continuously from real-world interactions. Field AI develops foundation models for unstructured environments, allowing robots to operate without GPS or pre-defined maps in settings like construction sites and oil fields.
Learning in Simulation#
Training a VLA model requires data: millions of examples pairing visual inputs and language commands with correct actions. Collecting this data in the physical world is slow, expensive, and limited by the number of robots available. A single robot can only attempt so many grasps per day. Failures risk damaging hardware. Certain scenarios, like recovering from falls or handling dangerous materials, cannot be practiced safely.
Simulation offers an alternative. In virtual environments, robots can attempt tasks millions of times without physical constraints. Time accelerates. Failures cost nothing. Dangerous scenarios become routine training exercises. A simulated robot can practice catching falling objects, navigating through crowds, or manipulating fragile materials, all without risking real hardware or real people.
NVIDIA has built the most comprehensive simulation platform for humanoid robotics. Isaac Sim provides physics-accurate virtual environments where robots can learn locomotion, manipulation, and navigation. Omniverse connects Isaac Sim to a broader ecosystem of 3D tools, allowing engineers to create realistic virtual factories, homes, and outdoor environments. Project GR00T uses these platforms to train foundation models for humanoid behavior, with the goal of creating general-purpose AI that transfers across different robot bodies.
The critical challenge is the simulation-to-real gap, the difference between virtual physics and actual physics that causes behaviors learned in simulation to fail in the real world. A simulated cup has perfectly predictable properties. A real cup might be chipped, wet, or positioned slightly differently than any cup in the training data. Simulated lighting is consistent. Real lighting varies with time of day, weather, and reflections from nearby surfaces.
Engineers have developed techniques to bridge this gap. Domain randomization varies simulation parameters randomly during training: changing object sizes, surface textures, lighting conditions, and physics properties. A model trained across thousands of randomized variations learns to handle uncertainty, making it more robust when deployed on physical hardware. The approach has proven effective for manipulation tasks, where models trained with aggressive domain randomization transfer successfully to real robots.
Another approach combines simulation with limited real-world data. Rather than training entirely in simulation, engineers collect a smaller dataset of real robot behavior and use it to fine-tune models pre-trained in virtual environments. This hybrid approach captures the scale advantages of simulation while grounding the model in actual physics.
Tesla claims to generate massive amounts of synthetic training data through simulation, supplementing real-world data from Optimus prototypes working in its factories. The company's experience with synthetic data for autonomous vehicles informs its approach to robotics, though details remain proprietary.
The simulation-to-real gap has not been fully solved. Complex contact dynamics, soft materials, and fluid interactions remain difficult to simulate accurately. Tasks involving cloth, liquids, or human bodies push current simulators beyond their reliable range. Progress continues, but humanoid robots in 2025 still require real-world fine-tuning to perform reliably outside controlled environments.
The Integration Challenge#
Hardware and software must work together within constraints that neither can escape. A model that requires 100 milliseconds to process a camera frame cannot support reflexive grasping. A chip that consumes 500 watts cannot operate in a battery-powered humanoid. And as we saw in examining the humanoid body, most robots today operate for only about two hours on a single charge, a constraint that shapes every decision about compute power. Running large AI models demands energy; the most capable models may be too power-hungry for mobile deployment. The best algorithm running on inadequate hardware will fail, as will the most powerful chip running poorly designed software.
This integration challenge is driving different strategic approaches. Tesla pursues full vertical integration, designing its own chips, training its own models, and building its own robots. This Apple-like approach maximizes control but requires excellence across every domain. Tesla is developing its Hardware 5 AI chip specifically for Optimus, building on its Dojo architecture to optimize inference for humanoid tasks.
Most competitors follow a more modular strategy. Figure AI uses NVIDIA hardware but develops its own Helix software. Sanctuary AI combines off-the-shelf sensors with proprietary AI. This approach allows companies to focus on their strengths while leveraging others' advances in complementary areas. The risk is dependence: if NVIDIA changes its platform or pricing, companies built on its ecosystem must adapt.
Google represents a third path. Its Tensor Processing Units excel at AI workloads, and DeepMind's robotics research pushes the boundaries of learned behavior. Partnerships with companies like Apptronik could see Google's AI running on third-party hardware, creating a potential "Android for robots" model where Google provides intelligence and partners provide bodies.
TSMC, the world's leading semiconductor foundry, benefits regardless of which architectural approach wins. Its advanced manufacturing processes, including 3nm and 2nm nodes, produce chips for NVIDIA, Qualcomm, and many startups. As humanoid production scales, TSMC's fabrication capacity could become a critical bottleneck and a source of significant value capture.
The Mind Emerges from the Loop#
The perception-action loop that enables a robot to grasp a coffee cup is not, in itself, a mind. It is a mechanism, sophisticated but fundamentally reactive. The robot does not want the cup. It does not know what coffee is. It executes learned behaviors in response to sensory inputs, with no inner experience we can detect or measure.
Yet as these systems grow more capable, the question Arthur C. Clarke raised becomes harder to dismiss. When a robot can follow complex instructions, adapt to novel situations, explain its reasoning in natural language, and learn from its mistakes, at what point does processing become something more? The question may be unanswerable, or it may be the wrong question to ask. What matters for the humans living and working alongside these machines is not whether they have minds in some philosophical sense, but whether they behave as if they do.
The humanoid brain in 2025 is a work in progress. No architecture has proven dominant. No company has achieved artificial general intelligence for robotics. The models are impressive but brittle, capable of remarkable feats in some contexts while failing at tasks a child could accomplish. The hardware is powerful but power-hungry, fast but not fast enough.
What has changed is the trajectory. Vision-language-action models have unified capabilities that were separate for decades. Simulation has unlocked training at scales impossible in the physical world. Hardware continues its exponential improvement. The gap between current systems and human-level embodied intelligence remains vast, but it is shrinking in ways that were not predictable even five years ago.
The body is ready. The hands can grasp. Now the question is whether the brain can learn to use them, and what it will mean for us when it does.
References: