paper-with-me

홈 › Papers

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

2026-05-28 · Jusuk Lee, Seungjae Lee, Jonghun Shin, Hoseong Jung, Sungha Kim, Daesol Cho, H. Jin Kim, Jia-Bin Huang, Furong Huang arxiv

Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recognition or vision-language alignment, leaving motion understanding to downstream policies. We introduce DynaFLIP, a dynamics-aware multimodal pre-training framework that pushes motion understanding upstream into perception. We construct image-language-3D flow triplets from heterogeneous human and robot videos, and use these triplets as training-time supervision to shape an image-only encoder. Our key idea is to encourage the three modalities to span a small simplex volume in the shared hyperspherical space -- a smaller simplex volume indicating stronger alignment. To avoid the geometric ambiguity and trivial collapse of naive volume minimization, we combine simplex-volume minimization with a cosine regularizer and a contrastive objective. Our analyses show that DynaFLIP focuses on control-relevant regions critical for manipulation. The resulting dynamics-aware representations serve as reusable visual backbones and consistently outperform baselines across diverse downstream policies, including VLAs. We validate this across diverse simulation and real-world setups, with gains reaching +22.5% under out-of-distribution scenarios. Our results suggest that robot generalization improves when visual representations are trained to encode not just what is present, but how the world changes under action.

📄 PDF Abstract BibTeX arXiv:2605.30350

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

Learning to Optimize Edge Robotics: A Fast Integrated Perception-Motion-Communication Approach

2025-10-18 · Dan Guo, Xibin Jin, Shuai Wang, Zhigang Wen 외 arxiv

Edge robotics involves frequent exchanges of large-volume multi-modal data. Existing methods ignore the interdependency between robotic functionalities and communication conditions, leading to excessive communication ove…

Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought

2026-03-24 · Yunheng Li, Hangyi Kuang, Hengrui Zhang, Jiangxia Cao 외 arxiv

Multimodal Chain-of-Thought (CoT) reasoning requires large vision-language models to construct reasoning trajectories that interleave perceptual grounding with multi-step inference. However, existing Reinforcement Learni…

Reinforcement LearningMultimodal ReasoningVisual Grounding

Multi-modal Visual Place Recognition in Dynamics-Invariant Perception Space

2021-05-17 · Lin Wu, Teng Wang, Changyin Sun

Visual place recognition is one of the essential and challenging problems in the fields of robotics. In this letter, we for the first time explore the use of multi-modal fusion of semantic and visual modalities in dynami…

SegmentationSemantic SegmentationVisual Place Recognition

Cross-Modal Benchmarking for Robotic Perception in Natural Environments

2026-06-10 · David Hall, Joshua Knights, Mark Cox, Peyman Moghadam arxiv

Natural environments present a complex challenge to robotics perception systems. Current models, particularly vision foundation models, are largely trained on structured, urban environments leading to weaknesses in their…

Depth Estimation

Align then Adapt: Rethinking Parameter-Efficient Transfer Learning in 4D Perception

2026-02-26 · Yiding Sun, Jihua Zhu, Haozhe Cheng, Chaoyi Lu 외 arxiv

Point cloud video understanding is critical for robotics as it accurately encodes motion and scene interaction. We recognize that 4D datasets are far scarcer than 3D ones, which hampers the scalability of self-supervised…

3D Action RecognitionSemantic SegmentationAction SegmentationTransfer Learning