paper-with-me

Papers

OmniVLA-RL: A Vision-Language-Action Model with Spatial Understanding and Online RL

2026-04-20 · Haoxiang Jie, Yaoyuan Yan, Xiangyu Wei, Kailin Wang, Hongjie Yan, Zhiyou Heng, Daocheng Chen arxiv

Visual-Language-Action (VLA) models represent a paradigm shift in embodied AI, yet existing frameworks often struggle with imprecise spatial perception, suboptimal multimodal fusion, and instability in reinforcement learning. To bridge these gaps, we propose OmniVLA-RL, a novel architecture that leverages a Mix-of-Transformers (MoT) design to synergistically integrate reasoning, spatial, and action experts. Furthermore, we introduce Flow-GSPO, which reformulates flow matching as a Stochastic Differential Equation (SDE) process and integrates it with Group Segmented Policy Optimization (GSPO) to enhance action precision and training robustness. Extensive evaluations on the LIBERO and LIBERO-Plus benchmarks demonstrate that OmniVLA-RL achieves decent overall performance and surpasses mainstream existing methods, effectively overcoming the fundamental limitations of current VLA models.

📄 PDF Abstract BibTeX arXiv:2604.17706

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation

2025-11-03 · Heyu Guo, Shanmu Wang, Ruichun Ma, Shiqi Jiang 외 arxiv

Vision-language-action (VLA) models have shown strong generalization for robotic action prediction through large-scale vision-language pretraining. However, most existing models rely solely on RGB cameras, limiting their…

OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation

2025-09-23 · Noriaki Hirose, Catherine Glossop, Dhruv Shah, Sergey Levine arxiv

Humans can flexibly interpret and compose different goal specifications, such as language instructions, spatial coordinates, or visual references, when navigating to a destination. In contrast, most existing robotic navi…

Robot Navigation

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies

2026-07-06 · Adrian Szvoren, Dimitrios Kanoulas, Nilufer Tuptuk arxiv

Vision-language-action (VLA) models enable robot navigation from natural language and visual goals, but remain susceptible to perceptual distractions and ambiguous scene interpretations. This paper presents the first emp…

Semantic SegmentationRobot NavigationVisual Grounding

TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding

2026-02-23 · Fan Yang, Shurong Zheng, Hongyin Zhao, Yufei Zhan 외 arxiv

Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, current approaches focus predominantly on global image understanding, strug…

Trajectory PredictionScene UnderstandingLogical Reasoning

Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding

2025-07-01 · Tao Lin, Gen Li, Yilei Zhong, Yanwen Zou 외 arxiv

Vision-Language-Action (VLA) models have emerged as a promising framework for enabling generalist robots capable of perceiving, reasoning, and acting in the real world. These models usually build upon pretrained Vision-L…

Depth EstimationPoint Clouds