paper-with-me

홈 › Papers

TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers

2026-01-20 · Bin Yu, Shijie Lian, Xiaopeng Lin, Yuliang Wei, Zhaolong Shen, Changti Wu, Yuzhuo Miao, Xinming Wang, Bailing Wang, Cong Huang, Kai Chen arxiv

The fundamental premise of Vision-Language-Action (VLA) models is to harness the extensive general capabilities of pre-trained Vision-Language Models (VLMs) for generalized embodied intelligence. However, standard robotic fine-tuning inevitably disrupts the pre-trained feature space, leading to "catastrophic forgetting" that compromises the general visual understanding we aim to leverage. To effectively utilize the uncorrupted general capabilities of VLMs for robotic tasks, we propose TwinBrainVLA, which coordinates two isomorphic VLM pathways: a frozen generalist (also called "Left Brain") and a trainable specialist (also called "Right Brain"). Our architecture utilizes a Asymmetric Mixture-of-Transformers (AsyMoT) mechanism, enabling the Right Brain to dynamically query and fuse intact semantic knowledge from the Left Brain with proprioceptive states. This fused representation conditions a flow-matching action expert for precise continuous control. Empirical results on SimplerEnv and RoboCasa benchmarks demonstrate that by explicitly retaining general capabilities, TwinBrainVLA achieves substantial performance gains over baseline models in complex manipulation tasks.

📄 PDF Abstract BibTeX arXiv:2601.14133

Code (0)

등록된 구현이 없습니다.

Tasks

Continuous Control

Similar Papers 제목 키워드 기반

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

2026-08-31 · Shaoan Wang, Aocheng Luo, Fei Huang, Jingyi Xu 외 hf

Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial pri…

Zero-shot GeneralizationReinforcement LearningInstruction FollowingSpatial Reasoning

Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision

2026-01-27 · Zhixiang Wei, Yi Li, Zhehan Kan, Xinghua Jiang 외 arxiv

Despite the significant advancements represented by Vision-Language Models (VLMs), current architectures often exhibit limitations in retaining fine-grained visual information, leading to coarse-grained multimodal compre…

EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models

2024-06-09 · Mengfei Du, Binhao Wu, Zejun Li, Xuanjing Huang 외

The recent rapid development of Large Vision-Language Models (LVLMs) has indicated their potential for embodied tasks.However, the critical skill of spatial understanding in embodied environments has not been thoroughly …

Benchmarking

GenRL: Multimodal-foundation world models for generalization in embodied agents

2024-06-26 · Pietro Mazzaglia, Tim Verbelen, Bart Dhoedt, Aaron Courville 외

Learning generalist embodied agents, able to solve multitudes of tasks in different domains is a long-standing problem. Reinforcement learning (RL) is hard to scale up as it requires a complex reward design for each task…

BenchmarkingReinforcement Learning (RL)

UIPro: Unleashing Superior Interaction Capability For GUI Agents

2025-09-22 · Hongxin Li, Jingran Su, Jingfan Chen, Zheng Ju 외 arxiv

Building autonomous agents that perceive and operate graphical user interfaces (GUIs) like humans has long been a vision in the field of artificial intelligence. Central to these agents is the capability for GUI interact…