• 5 mins read
  • Published

Dyna Robotics Tests DYNA-2 Robot Model Trained on Human Video

Noel Sharkey Technology, AI and robotics editor Science.Report

Post by Noel Sharkey

Dyna Robotics Tests DYNA-2 Robot Model Trained on Human Video Science.Report © science.report
Dyna Robotics Tests DYNA-2 Robot Model Trained on Human Video © science.report

Dyna Robotics has introduced DYNA-2, a robot foundation model trained on over one million hours of human egocentric video, aiming to improve robot learning for physical tasks without relying on robot action data

Dyna Robotics has released DYNA-2, a robot foundation model trained exclusively on more than one million hours of human egocentric video. The company, based in Redwood City, California, claims this approach allows robots to acquire physical skills by observing how humans interact with objects and environments, rather than relying on robot-generated action data. The dataset, representing approximately 170 years of continuous waking human activity, is intended to address the challenge of scaling robot learning beyond the limits of manually collected teleoperation data.

According to Dyna Robotics, DYNA-2 employs a world-modeling architecture that predicts both the next frame and the next action in a sequence. By learning from human video, the model is designed to generalize across different robot hardware, including stationary arms, humanoid prototypes, and dexterous robotic hands. The company reports that DYNA-2 was trained without exposure to robot action data during pretraining, but can be adapted to new platforms with a small amount of local fine-tuning.

Evaluation and Measured Performance

In internal tests, Dyna Robotics states that DYNA-2 increased task success rates in high-precision manufacturing from 20% to between 80% and 90% when compared to models trained on less data. The company also reports that, in one demonstration, 13 minutes of fine-tuning data enabled DYNA-2 to control a pair of five-fingered robotic hands to open a bottle cap. Across 15 benchmark tasks, models trained with larger volumes of human video consistently outperformed those with less exposure. In a zero-shot customer deployment, DYNA-2 reportedly achieved an 87% quality pass rate, compared to 46% for the earlier DYNA-1 model, which used a vision-language-action architecture.

DYNA-2 was also tested for resilience to physical disturbances. During tasks such as chopping food and clearing workspaces, the model was able to recover from disruptions without human intervention, while the previous version required manual resets. The company attributes a 133% improvement in instruction-following tasks to its video co-training method, which enables robots to perform varied physical motions in response to commands.

Training Data and Transfer

The DYNA-2 model was trained entirely on human egocentric video, with no robot action data included during pretraining. This strategy is intended to make robot learning more scalable as capabilities expand, reducing the need for labor-intensive teleoperation data collection. Dyna Robotics claims that knowledge acquired from human video can transfer across different robot platforms, with adaptation requiring only a few hours of additional data. The company's approach contrasts with previous methods that depend heavily on robot-specific demonstrations.

Other research teams have explored similar strategies for transferring human skills to robots. For example, efforts to enable robots to adapt to human partners during shared physical tasks, such as those described in recent studies on collaborative lifting robots, highlight the broader trend toward leveraging human data for robot learning. However, the scale and exclusive reliance on human video in DYNA-2's training set distinguish it from most prior work.

Deployment and Remaining Questions

Dyna Robotics reports that robots powered by its earlier DYNA-1 model are already deployed in hotels, restaurants, and laundromats. The company positions DYNA-2 as a step toward enabling robots to learn new physical tasks without extensive robot-specific training data. However, the evidence for DYNA-2's capabilities comes primarily from company demonstrations and internal benchmarks. Independent verification of the model's performance, generalization, and safety in diverse real-world environments has not yet been reported.

Key limitations remain. The company has not disclosed the full details of its training dataset, including the diversity of environments, object types, or potential biases present in the human video. It is also unclear how DYNA-2 performs in unstructured or safety-critical settings, or how much human oversight is required during deployment. As with other foundation models, the transferability of skills learned from video to novel tasks and hardware may be constrained by differences between human and robot embodiment, sensor modalities, and actuation limits.

Dyna Robotics was founded by Lindon Gao, York Yang, and Jason Ma, a former DeepMind research scientist, and is backed by investors including CRV and First Round. The company has not announced plans for public release of the DYNA-2 model or its training data.

Foundation models in robotics are large-scale machine learning systems trained on broad datasets intended to support a wide range of downstream tasks. Unlike task-specific models, foundation models are designed to generalize across different robots, environments, and objectives. Training on human egocentric video allows these models to capture patterns of physical interaction, but transferring this knowledge to robots requires careful adaptation due to differences in hardware and sensing. The effectiveness of such transfer depends on the similarity between the training data and the deployment context, as well as the ability to fine-tune the model with limited robot-specific data. Ongoing research aims to clarify the limits of generalization, safety, and reliability for foundation models in real-world robotics applications.

Related articles