paper-with-me

Papers

Vision-Language-Action Models: Concepts, Progress, Applications and Challenges

2025-05-07 · Ranjan Sapkota, Yang Cao, Konstantinos I. Roumeliotis, Manoj Karkee

Vision-Language-Action (VLA) models mark a transformative advancement in artificial intelligence, aiming to unify perception, natural language understanding, and embodied action within a single computational framework. This foundational review presents a comprehensive synthesis of recent advancements in Vision-Language-Action models, systematically organized across five thematic pillars that structure the landscape of this rapidly evolving field. We begin by establishing the conceptual foundations of VLA systems, tracing their evolution from cross-modal learning architectures to generalist agents that tightly integrate vision-language models (VLMs), action planners, and hierarchical controllers. Our methodology adopts a rigorous literature review framework, covering over 80 VLA models published in the past three years. Key progress areas include architectural innovations, parameter-efficient training strategies, and real-time inference accelerations. We explore diverse application domains such as humanoid robotics, autonomous vehicles, medical and industrial robotics, precision agriculture, and augmented reality navigation. The review further addresses major challenges across real-time control, multimodal action representation, system scalability, generalization to unseen tasks, and ethical deployment risks. Drawing from the state-of-the-art, we propose targeted solutions including agentic AI adaptation, cross-embodiment generalization, and unified neuro-symbolic planning. In our forward-looking discussion, we outline a future roadmap where VLA models, VLMs, and agentic AI converge to power socially aligned, adaptive, and general-purpose embodied agents. This work serves as a foundational reference for advancing intelligent, real-world robotics and artificial general intelligence. >Vision-language-action, Agentic AI, AI Agents, Vision-language Models

📄 PDF Abstract BibTeX arXiv:2505.04769

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous VehiclesNatural Language UnderstandingVision-Language-Action

Similar Papers 제목 키워드 기반

Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models

2024-12-11 · Quang-Hung Le, Long Hoang Dang, Ngan Le, Truyen Tran 외

Existing Large Vision-Language Models (LVLMs) excel at matching concepts across multi-modal inputs but struggle with compositional concepts and high-level relationships between entities. This paper introduces Progressive…

Question AnsweringVisual GroundingVisual Reasoning

Enhancing Zero-Shot Image Recognition in Vision-Language Models through Human-like Concept Guidance

2025-03-20 · Hui Liu, Wenya Wang, Kecheng Chen, Jie Liu 외

In zero-shot image recognition tasks, humans demonstrate remarkable flexibility in classifying unseen categories by composing known simpler concepts. However, existing vision-language models (VLMs), despite achieving sig…

Prompt EngineeringZero-shot Generalization

NEUCORE: Neural Concept Reasoning for Composed Image Retrieval

2023-10-02 · Shu Zhao, Huijuan Xu

Composed image retrieval which combines a reference image and a text modifier to identify the desired target image is a challenging task, and requires the model to comprehend both vision and language modalities and their…

Concept AlignmentImage RetrievalMultiple Instance LearningRetrieval+1

What value do explicit high level concepts have in vision to language problems?

2015-06-03 · CVPR 2016 6 · Qi Wu, Chunhua Shen, Lingqiao Liu, Anthony Dick 외

Much of the recent progress in Vision-to-Language (V2L) problems has been achieved through a combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). This approach does not explicitly rep…

Image CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Solving Bongard Problems with a Visual Language and Pragmatic Reasoning

2018-04-12 · Stefan Depeweg, Constantin A. Rothkopf, Frank Jäkel

More than 50 years ago Bongard introduced 100 visual concept learning problems as a testbed for intelligent vision systems. These problems are now known as Bongard problems. Although they are well known in the cognitive …

Bayesian Inference