paper-with-me

홈 › Papers

Foundation Model Driven Robotics: A Comprehensive Review

2025-07-14 · Muhammad Tayyab Khan, Ammar Waheed arxiv

The rapid emergence of foundation models, particularly Large Language Models (LLMs) and Vision-Language Models (VLMs), has introduced a transformative paradigm in robotics. These models offer powerful capabilities in semantic understanding, high-level reasoning, and cross-modal generalization, enabling significant advances in perception, planning, control, and human-robot interaction. This critical review provides a structured synthesis of recent developments, categorizing applications across simulation-driven design, open-world execution, sim-to-real transfer, and adaptable robotics. Unlike existing surveys that emphasize isolated capabilities, this work highlights integrated, system-level strategies and evaluates their practical feasibility in real-world environments. Key enabling trends such as procedural scene generation, policy generalization, and multimodal reasoning are discussed alongside core bottlenecks, including limited embodiment, lack of multimodal data, safety risks, and computational constraints. Through this lens, this paper identifies both the architectural strengths and critical limitations of foundation model-based robotics, highlighting open challenges in real-time operation, grounding, resilience, and trust. The review concludes with a roadmap for future research aimed at bridging semantic reasoning and physical intelligence through more robust, interpretable, and embodied models.

📄 PDF Abstract BibTeX arXiv:2507.10087

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningScene Generation

Similar Papers 제목 키워드 기반

Robot Learning from Human Videos: A Survey

2026-04-30 · Junyi Ma, Erhang Zhang, Haoran Yang, Ditao Li 외 arxiv

A critical bottleneck hindering further advancement in embodied AI and robotics is the challenge of scaling robot data. To address this, the field of learning robot manipulation skills from human video data has attracted…

Robot ManipulationVideo Generation

Large language model-based task planning for service robots: A review

2025-10-27 · Shaohan Bian, Ying Zhang, Guohui Tian, Zhiqiang Miao 외 arxiv

With the rapid advancement of large language models (LLMs) and robotics, service robots are increasingly becoming an integral part of daily life, offering a wide range of services in complex environments. To deliver thes…

Prompt Engineering

Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions

2026-04-16 · Aggelos Psiris, Vasileios Argyriou, Evangelos K. Markakis, Panagiotis Sarigiannidis 외 arxiv

Over the recent years, the field of robotics has been undergoing a transformative paradigm shift from fixed, single-task, domain-specific solutions towards adaptive, multi-function, generalpurpose agents, capable of oper…

A Review of Generative AI in Aquaculture: Foundations, Applications, and Future Directions for Smart and Sustainable Farming

2025-07-16 · Waseem Akram, Muhayy Ud Din, Lyes Saad Soud, Irfan Hussain arxiv

Generative Artificial Intelligence (GAI) has rapidly emerged as a transformative force in aquaculture, enabling intelligent synthesis of multimodal data, including text, images, audio, and simulation outputs for smarter,…

Vision Language Action Models in Robotic Manipulation: A Systematic Review

2025-07-14 · Muhayy ud Din, Waseem Akram, Lyes Saad Saoud, Jan Rosell 외

Vision Language Action (VLA) models represent a transformative shift in robotics, with the aim of unifying visual perception, natural language understanding, and embodied control within a single learning framework. This …

Dataset GenerationNatural Language UnderstandingVision-Language-Action