Humanoid robots can walk and manipulate objects in demonstrations, but reliable factory work still depends on real-world data, simulation, measurable safety systems and continuous human oversight.
A humanoid robot that appears to roam a factory independently may in fact be operating inside a tightly controlled zone with safeguards and a teleoperator nearby. That distinction sits at the center of the industry's deployment problem: a convincing demonstration is not the same as a dependable worker.
Rick Balzano discussed that gap in Lexicon. As Vice President of Go-to-Market for AI Services, Safety and Robotics at TaskUs, he works with companies developing robotics, autonomous vehicles and other AI-powered systems. Much of that work concerns the less visible infrastructure behind physical automation: collecting, labeling and validating the data machines need before they can act around people and equipment.
For a scientific assessment, the evidence should be separated into several categories: controlled laboratory evaluations, physical trials, operational safety records and commercial forecasts. The material available here does not provide a peer-reviewed sample size, p-value, confidence interval or independently audited failure rate for general-purpose humanoid factory work. That limitation does not invalidate the engineering progress, but it does mean that marketing demonstrations should not be treated as population-level evidence of reliability.
The hidden control layer
Large language models mainly process digital inputs such as text and images. An embodied AI system must turn sensor data into physical behavior. It has to perceive objects and surfaces, select an action, apply an appropriate amount of force and recover when the environment does not match its expectations.
That requires data that cannot simply be downloaded from the internet. Developers need first-person video, sensor readings, demonstrations of physical tasks and records of how objects respond when touched, lifted or moved. The training material must also include varied lighting, layouts, objects and working practices because a machine that performs well in one narrow setting may fail when the surroundings change.
In robotics, this is a distribution-shift problem: the probability distribution encountered during deployment can differ sharply from the one used during training. Researchers at MIT and Stanford have studied related problems in robot learning, but results from a particular laboratory task cannot automatically establish safe performance in a warehouse or automotive plant. A credible deployment claim therefore needs task definitions, exposure conditions, intervention rates and repeatability measurements, not only a successful video.
The label humanoid can obscure the engineering trade-off. Balzano's assessment is that many current machines are still specialized robots presented in a human-like form. The body may help with tasks designed around human workplaces, but wheels, claws or dedicated tools could be more effective in other settings. The longer-term design goal is a capable control system that can work with different physical configurations rather than a single universal body.
What demonstrations leave out
Online video usually shows the successful portion of an experiment. It may not show remote assistance, preparation of the work area, restricted operating boundaries or the resets required after a failed movement. A robot can look as though it is freely navigating a factory while remaining confined to a defined area where humans and software are ready to intervene.
That does not make the progress fraudulent. Locomotion, perception and manipulation have improved. It does mean that appearance is weak evidence for general-purpose capability. A short demonstration can establish that a system completed a selected task under selected conditions; it cannot establish that the same system will repeat the task across different factories without support.
The hardest training examples are often the rare events that create the greatest danger. Developers cannot routinely make costly machines fall down stairs or collide with equipment merely to gather failure data. Yet a robot still needs a response for an unexpected obstruction, a loss of balance or an object that behaves differently from the training examples. At present humans provide much of the teaching that is sometimes described as self-learning.
Safety engineering must also account for hazards that are statistically rare but physically severe. This is why failure reporting should include exposure time, number of task cycles, near misses, emergency stops, human interventions and the severity of plausible outcomes. Without those denominators, a statement that a robot completed a task successfully says little about how often it failed or how difficult recovery was.
Simulation stops short of reality
Simulation is central because it allows companies to repeat movements at scale without damaging hardware. A virtual environment can generate synthetic hazards and edge cases while keeping the financial and physical cost of failure low. It also makes it possible to test behavior that would be dangerous or impractical to reproduce repeatedly in a real factory.
But simulation is an approximation rather than a guarantee. It may not reproduce every irregular surface, sensor error, mechanical fault, obstruction or unpredictable human movement. Balzano described the problem through a 90-10 rule: simulation may take a robot most of the way toward a task, while the final portion can consume a disproportionate share of the deployment budget.
The numerical picture is deliberately limited but revealing. A robot shown in a factory video might actually be restricted to a 40-by-40 area with safeguards and teleoperators. Balzano estimates that the transition to useful operation across varied real-world settings could take five to ten years. He also described simulation as potentially delivering 90 percent of the progress while leaving the final 10 percent as the expensive practical challenge. These figures are expert estimates, not the result of a published controlled study, and should be read as an illustration of deployment economics rather than a measured universal law.
Training in one virtual factory does not automatically prepare a machine for ten physical factories. Each site has different layouts, surfaces, lighting, equipment and routines. A system must therefore be tested in the environments where it will operate and supported by procedures for intervention and recovery.
Safety is the scaling constraint
Before humanoids can become routine workers in warehouses, logistics facilities or automotive plants, deployment requires more than a capable robot. Facilities need dependable connectivity, remote intervention, physical recovery tools and clear instructions for employees sharing the workspace. The machine must also encounter enough environmental variation during development to avoid overfitting to a narrow set of conditions.
That shift is visible in Agility Robotics' partnership with FORT Robotics. The companies announced a memorandum of understanding on October 1, 2026, covering joint work on hardware, compliance and deployment support for Digit 5 in warehouses, manufacturing and logistics. Agility describes the platform as being engineered for cooperatively safe work at scale, with safe human detection, safety cues and safe motion control. The emphasis is significant: the commercial bottleneck is increasingly a demonstrable safety architecture rather than the robot's human-like silhouette. The partnership is described in an Agility safety announcement.
The regulatory picture remains incomplete. ISO 25785-1, discussed as a prospective international standard for dynamically stable robots including humanoids, was still at the committee-draft stage on October 2, 2026. It therefore should not be presented as a finalized mandatory rule. Until standards mature, manufacturers and buyers must combine existing machinery-safety practice, site-specific risk assessments, documented validation and transparent limits.
Balzano expects automotive assembly, warehousing and logistics to provide some of the earliest convincing financial returns because these sectors already use automation and tend to offer more structured environments. Yet productivity and purchase price will not determine whether humanoids scale by themselves. Reliability and integration determine whether a machine can contribute revenue; safety determines whether companies can use it broadly around people.
Industrial plans illustrate both the opportunity and the uncertainty. Hyundai and Boston Dynamics have opened a robotics center at Metaplant America in Georgia as a test bed and training center for Atlas integration. The center's early focus is repetitive parts sequencing and heavy-lifting tasks, while component assembly is planned for 2030. Hyundai has also described an ambition to deploy 25,000 Atlas units across its factories worldwide in the coming years. These are deployment plans and operational milestones, not independent evidence that every unit will achieve human-level generality. The facility's role is summarized in an factory testing report.
In China, Agibot has characterized 2026 and 2027 as critical years for large-scale factory adoption, with expansion plans spanning consumer electronics, semiconductor packaging and testing, automobiles, logistics and warehousing. Such plans show where industrial demand is being directed, but planned scale should not be confused with validated safety performance. The relevant question is how many operating hours, task cycles and interventions will be documented after deployment.
This is also a data-governance problem inside the factory. Someone must define which behavior counts as safe, record failures, validate labels and decide when a system is ready to move from controlled testing into live work. If human intervention remains essential but is hidden behind polished footage, buyers and workers cannot accurately judge the system's operating limits.
The same distinction matters when comparing humanoids with other robots. A human-shaped machine is not automatically more flexible than a wheeled platform or a purpose-built industrial arm. Its value depends on whether the hardware and control system can perform a useful task repeatedly at an acceptable level of risk, not on whether its silhouette resembles a worker.
The issue has a wider robotics context. NASA's work with human-supervised machines for lunar and Mars tasks, described in an earlier robotics report, illustrates the same engineering principle: supervision can extend what a machine does, but it also defines the boundary between assisted operation and independence.
Evidence before adoption
Balzano's account shifts attention away from viral clips and toward distributional coverage: whether training and testing include enough different places, objects, lighting conditions and human behaviors. Without that coverage a model may work reliably only in the environment that shaped it. A factory trial is therefore not merely a product demonstration; it is an examination of how the system behaves when reality departs from its preparation.
Peer-reviewed venues such as Nature and Science routinely distinguish between a proof-of-concept result and a validated performance claim. The same discipline is needed here. A serious report should state the robot model, task definition, number of trials, operating hours, environmental conditions, intervention frequency, failure taxonomy and uncertainty around each estimate. If a company does not publish those details, the appropriate conclusion is that the evidence remains incomplete, not that the system has failed or succeeded universally.
The central question is not whether humanoid robots can walk, grasp or complete a carefully selected action. Some can. The practical question is whether they can perform useful work repeatedly while managing rare failures, adapting to a new site and allowing humans to intervene before a mistake becomes damage or injury.
Simulation and real-world testing should be treated as complementary evidence rather than interchangeable proof. Simulation expands the number of situations a robot can encounter safely, while physical trials expose hardware limits, sensor noise and human unpredictability. A robot trained in simulation has learned within a designed approximation of the world; deployment begins when that approximation is no longer enough.
That is why the industry's most credible path is slower and less theatrical than its marketing videos. Humanoid robots have demonstrated meaningful advances, but they have not yet demonstrated general-purpose factory independence on the evidence presented here. Until real-world data, human oversight, compliance work and safety procedures are treated as core parts of the product rather than backstage support, the responsible description is specialized automation in a human-shaped body-not a ready-made replacement for a worker.