paper-with-me

홈 › Papers

Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions

2026-04-16 · Aggelos Psiris, Vasileios Argyriou, Evangelos K. Markakis, Panagiotis Sarigiannidis, Efstratios Gavves, Kostas Bekris, Arash Ajoudani adn Georgios Th. Papadopoulos arxiv

Over the recent years, the field of robotics has been undergoing a transformative paradigm shift from fixed, single-task, domain-specific solutions towards adaptive, multi-function, generalpurpose agents, capable of operating in complex, open-world, and dynamic environments. This tremendous advancement is primarily driven by the emergence of Foundation Models (FMs), i.e., large-scale neural-network architectures trained on massive, heterogeneous datasets that provide unprecedented capabilities in multi-modal understanding and reasoning, long-horizon planning, and cross-embodiment generalization. In this context, the current study provides a holistic, systematic, and in-depth review of the research landscape of FMs in robotics. In particular, the evolution of the field is initially delineated through five distinct research phases, spanning from the early incorporation of Natural Language Processing (NLP) and Computer Vision (CV) models to the current frontier of multi-sensory generalization and real-world deployment. Subsequently, a highly-granular taxonomic investigation of the literature is performed, examining the following key aspects: a) the employed FM types, including LLMs, VFMs, VLMs, and VLAs, b) the underlying neural-network architectures, c) the adopted learning paradigms, d) the different learning stages of knowledge incorporation, e) the major robotic tasks, and f) the main real-world application domains. For each aspect, comparative analysis and critical insights are provided. Moreover, a report on the publicly available datasets used for model training and evaluation across the considered robotic tasks is included. Furthermore, a hierarchical discussion on the current open challenges and promising future research directions in the field is incorporated.

📄 PDF Abstract BibTeX arXiv:2604.15395

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pure Vision Language Action (VLA) Models: A Comprehensive Survey

2025-09-23 · Dapeng Zhang, Jing Sun, Chenghui Hu, Xiaoyan Wu 외 arxiv

The emergence of Vision Language Action (VLA) models marks a paradigm shift from traditional policy-based control to generalized robotics, reframing Vision Language Models (VLMs) from passive sequence generators into act…

Foundation Model Driven Robotics: A Comprehensive Review

2025-07-14 · Muhammad Tayyab Khan, Ammar Waheed arxiv

The rapid emergence of foundation models, particularly Large Language Models (LLMs) and Vision-Language Models (VLMs), has introduced a transformative paradigm in robotics. These models offer powerful capabilities in sem…

Multimodal ReasoningScene Generation

Robot Learning from Human Videos: A Survey

2026-04-30 · Junyi Ma, Erhang Zhang, Haoran Yang, Ditao Li 외 arxiv

A critical bottleneck hindering further advancement in embodied AI and robotics is the challenge of scaling robot data. To address this, the field of learning robot manipulation skills from human video data has attracted…

Robot ManipulationVideo Generation

Camera Trajectory Generation: A Comprehensive Survey of Methods, Metrics, and Future Directions

2025-06-01 · Zahra Dehghanian, Pouya Ardekhani, Amir Vahedi, Hamid Beigy 외

Camera trajectory generation is a cornerstone in computer graphics, robotics, virtual reality, and cinematography, enabling seamless and adaptive camera movements that enhance visual storytelling and immersive experience…

Visual Storytelling

Vision-Language-Action Models: Concepts, Progress, Applications and Challenges

2025-05-07 · Ranjan Sapkota, Yang Cao, Konstantinos I. Roumeliotis, Manoj Karkee

Vision-Language-Action (VLA) models mark a transformative advancement in artificial intelligence, aiming to unify perception, natural language understanding, and embodied action within a single computational framework. T…

Autonomous VehiclesNatural Language UnderstandingVision-Language-Action