• 5 mins read
  • Published

Figure and Nscale plan massive GPU buildout for robot AI training

Noel Sharkey Technology, AI and robotics editor Science.Report

Post by Noel Sharkey

Figure and Nscale plan massive GPU buildout for robot AI training Science.Report © science.report
Figure and Nscale plan massive GPU buildout for robot AI training © science.report

Figure and Nscale have announced a multi-billion dollar partnership to deploy up to 100,000 NVIDIA GPUs for training general-purpose humanoid robot AI models, aiming to overcome the data and compute bottlenecks in physical intelligence research

Figure and Nscale are moving ahead with plans to build one of the largest AI computing setups to date, aiming to deploy up to 100,000 GPUs for training humanoid robot models. The deal, valued at $3.5 billion initially and possibly exceeding $6 billion, marks a major push to develop general-purpose physical intelligence. Both companies are betting that access to this much computing power will help them build robots that are more capable and adaptable. The planned scale is on par with the biggest AI infrastructure projects worldwide, including those at MIT and Stanford, which have long set the standard for computational resources in robotics and machine learning.

Figure's goal is to move past controlled demos and build robots that can handle a wide range of real-world tasks. Its Helix AI system, designed for humanoid robots, is being trained on increasingly large and varied datasets of human activity. As the data grows, so does the need for specialized hardware. Nscale plans to deploy NVIDIA Vera Rubin platform GPUs, starting with a major installation in Barstow, Texas, targeted for the second half of 2027. This infrastructure is intended to support not just model training, but also simulation and deployment, using NVIDIA's robotics simulation tools to connect virtual learning with real-world performance. The full timeline for all 100,000 GPUs hasn't been shared, and the project is still in the planning stage, not yet an operational cluster. This cautious approach is similar to how large scientific collaborations like CERN and NASA roll out new infrastructure.

Figure has changed how it collects training data, moving away from outside suppliers and building its own pipeline, called Index. Index gathers and filters video data of human activity from contributors around the world. Figure says Index has passed 264,000 downloads in 108 countries, with over 44,000 weekly active users and enough new video uploaded to generate 30 minutes of data every second. The system automatically checks for technical, visual, and semantic quality, removes duplicates, and rebalances the dataset to cover a wide range of tasks and environments. Human reviewers check samples for fraud and other issues before the data is used for training. So far, Figure reports paying $15 million to contributors and plans to increase both data and compute spending as the project grows. This approach to curating and checking datasets follows best practices highlighted in peer-reviewed studies in Nature, where reproducibility and dataset diversity are seen as essential for strong AI development.

The partnership also includes a strategic investment from Nscale into Figure, and the companies are looking at ways to use humanoid robots in Nscale's own supply chain. While the main focus is on training and deploying general-purpose models, neither company has released specific benchmarks or independent evaluations of Helix's current abilities. The agreement is structured as a long-term infrastructure buildout, with the first deployments still years away and full capacity depending on future hardware delivery and integration. This phased approach is similar to other large AI infrastructure projects, such as those managed by the Max Planck Society, where gradual rollouts allow for ongoing testing and risk management.

Figure's planned GPU deployment is much larger than most current robotics research efforts. Earlier projects, such as modular platforms, have focused more on hardware flexibility than on compute scale. Figure is betting that massive, purpose-built AI infrastructure will speed up progress in physical intelligence, but there are still significant technical and operational risks. The company's use of proprietary data pipelines and simulation tools raises questions about reproducibility, safety, and whether results from virtual training will transfer to real-world settings. Recent reviews by the editorial board of Science have also pointed out the need for open benchmarks and independent validation in robotics and AI research.

So far, evidence for general-purpose robotic intelligence is limited to company demos and internal metrics. There have been no independent audits or peer-reviewed evaluations of Helix's performance on standard benchmarks. The partnership's success will depend not just on hardware delivery, but on whether simulated learning can be turned into reliable, safe, and adaptable behavior in the real world. Until independent results are published, claims of a breakthrough in humanoid robot capability should be viewed with caution. The real test will be whether these robots can perform outside the lab, under real-world conditions, and with proper human oversight.

Training AI models for humanoid robots requires huge amounts of data and computing power. Unlike language models, which learn from text, physical intelligence systems must process video, sensor data, and real-world interactions to generalize across tasks and environments. This calls for high-performance GPUs, strong data pipelines, simulation environments, and safety checks. The gap between simulated training and real-world deployment-the sim-to-real gap-remains a major challenge, as models that work well in virtual settings often struggle with the unpredictability of the physical world. Closing this gap is central to the promise and risk of large-scale robot AI projects like Figure's, and remains an active area of research at places like MIT and Stanford.

Related articles