D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI

2510.05684v1 cs.AI, cs.CV, cs.RO 2025-10-09

Авторы:

Suwhan Choi, Jaeyoon Jung, Haebin Seong, Minchan Kim, Minyeong Kim, Yongjun Cho, Yoonshik Kim, Yubeen Park, Youngjae Yu, Yunsung Lee

Abstract

Large language models leverage internet-scale text data, yet embodied AI remains constrained by the prohibitive costs of physical trajectory collection. Desktop environments -- particularly gaming -- offer a compelling alternative: they provide rich sensorimotor interactions at scale while maintaining the structured observation-action coupling essential for embodied learning. We present D2E (Desktop to Embodied AI), a framework that demonstrates desktop interactions can serve as an effective pretraining substrate for robotics embodied AI tasks. Unlike prior work that remained domain-specific (e.g., VPT for Minecraft) or kept data proprietary (e.g., SIMA), D2E establishes a complete pipeline from scalable desktop data collection to verified transfer in embodied domains. Our framework comprises three components: (1) the OWA Toolkit that unifies diverse desktop interactions into a standardized format with 152x compression, (2) the Generalist-IDM that achieves strong zero-shot generalization across unseen games through timestamp-based event prediction, enabling internet-scale pseudo-labeling, and (3) VAPT that transfers desktop-pretrained representations to physical manipulation and navigation. Using 1.3K+ hours of data (259 hours of human demonstrations, and 1K+ hours of pseudo-labeled gameplay), we achieve a total of 96.6% success rate on LIBERO manipulation and 83.3% on CANVAS navigation benchmarks. This validates that sensorimotor primitives in digital interactions exhibit sufficient invariance to transfer meaningfully to physical embodied tasks, establishing desktop pretraining as a practical paradigm for robotics. We will make all our work public, including the OWA toolkit, datasets of human-collected and pseudo-labeled, and VAPT-trained models available at https://worv-ai.github.io/d2e/

Ссылки и действия

Читать на arXiv Скачать PDF

Дополнительные ресурсы:

D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI

Авторы:

Abstract

Ссылки и действия

Связанные статьи

Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning

Unified World Models: Memory-Augmented Planning and Foresight for Visual Navigat...

The Safety Challenge of World Models for Embodied AI Agents: A Review

Robix: A Unified Model for Robot Interaction, Reasoning and Planning

sam-llm: interpretable lane change trajectoryprediction via parametric finetunin...

Навигация