paper-with-me

홈 › Papers

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding

2026-03-18 · Shuyao Shi, Kang G. Shin arxiv

Recent Multimodal Large Language Models (MLLMs) have shown high potential for spatial reasoning within 3D scenes. However, they typically rely on computationally expensive 3D representations like point clouds or reconstructed Bird's-Eye View (BEV) maps, or lack physical grounding to resolve ambiguities in scale and size. This paper significantly enhances MLLMs with egomotion modality data, captured by Inertial Measurement Units (IMUs) concurrently with the video. In particular, we propose a novel framework, called Motion-MLLM, introducing two key components: (1) a cascaded motion-visual keyframe filtering module that leverages both IMU data and visual features to efficiently select a sparse yet representative set of keyframes, and (2) an asymmetric cross-modal fusion module where motion tokens serve as intermediaries that channel egomotion cues and cross-frame visual context into the visual representation. By grounding visual content in physical egomotion trajectories, Motion-MLLM can reason about absolute scale and spatial relationships across the scene. Our extensive evaluation shows that Motion-MLLM makes significant improvements in various tasks related to 3D scene understanding and spatial reasoning. Compared to state-of-the-art (SOTA) methods based on video frames and explicit 3D data, Motion-MLLM achieves competitive accuracy while running $1.30\times$ and $1.61\times$ faster, respectively.

📄 PDF Abstract BibTeX arXiv:2603.17980

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingSpatial ReasoningPoint Clouds

Similar Papers 제목 키워드 기반

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent

2025-07-21 · Jiaao Li, Kaiyuan Li, Chen Gao, Yong Li 외

Egomotion videos are first-person recordings where the view changes continuously due to the agent's movement. As they serve as the primary visual input for embodied AI agents, making egomotion video reasoning more effici…

Multimodal Reasoning

Learning Spatial Common Sense with Geometry-Aware Recurrent Networks

2018-12-31 · CVPR 2019 6 · Hsiao-Yu Fish Tung, Ricson Cheng, Katerina Fragkiadaki

We integrate two powerful ideas, geometry and deep visual representation learning, into recurrent network architectures for mobile visual scene understanding. The proposed networks learn to "lift" and integrate 2D visual…

Common Sense ReasoningRepresentation LearningScene Understanding

Novel Diffusion Models for Multimodal 3D Hand Trajectory Prediction

2025-04-10 · Junyi Ma, Wentao Bao, Jingyi Xu, Guanzhong Sun 외

Predicting hand motion is critical for understanding human intentions and bridging the action space between human movements and robot manipulations. Existing hand trajectory prediction (HTP) methods forecast the future h…

DenoisingMambaTrajectory Prediction

CLiF-VQA: Enhancing Video Quality Assessment by Incorporating High-Level Semantic Information related to Human Feelings

2023-11-13 · Yachun Mi, Yu Li, Yan Shu, Chen Hui 외

Video Quality Assessment (VQA) aims to simulate the process of perceiving video quality by the human visual system (HVS). The judgments made by HVS are always influenced by human subjective feelings. However, most of the…

Video Quality AssessmentVisual Question Answering (VQA)

Egocentric Audio-Visual Object Localization

2023-03-23 · CVPR 2023 1 · Chao Huang, Yapeng Tian, Anurag Kumar, Chenliang Xu

Humans naturally perceive surrounding scenes by unifying sound and sight in a first-person view. Likewise, machines are advanced to approach human intelligence by learning with multisensory inputs from an egocentric pers…

ObjectObject Localization