paper-with-me

홈 › Papers

Egocentric Scene Understanding via Multimodal Spatial Rectifier

2022-07-14 · CVPR 2022 1 · Tien Do, Khiem Vuong, Hyun Soo Park

In this paper, we study a problem of egocentric scene understanding, i.e., predicting depths and surface normals from an egocentric image. Egocentric scene understanding poses unprecedented challenges: (1) due to large head movements, the images are taken from non-canonical viewpoints (i.e., tilted images) where existing models of geometry prediction do not apply; (2) dynamic foreground objects including hands constitute a large proportion of visual scenes. These challenges limit the performance of the existing models learned from large indoor datasets, such as ScanNet and NYUv2, which comprise predominantly upright images of static scenes. We present a multimodal spatial rectifier that stabilizes the egocentric images to a set of reference directions, which allows learning a coherent visual representation. Unlike unimodal spatial rectifier that often produces excessive perspective warp for egocentric images, the multimodal spatial rectifier learns from multiple directions that can minimize the impact of the perspective warp. To learn visual representations of the dynamic foreground objects, we present a new dataset called EDINA (Egocentric Depth on everyday INdoor Activities) that comprises more than 500K synchronized RGBD frames and gravity directions. Equipped with the multimodal spatial rectifier and the EDINA dataset, our proposed method on single-view depth and surface normal estimation significantly outperforms the baselines not only on our EDINA dataset, but also on other popular egocentric datasets, such as First Person Hand Action (FPHA) and EPIC-KITCHENS.

📄 PDF Abstract BibTeX arXiv:2207.07077

Code (1)

tien-d/EgoDepthNormal 공식 구현 pytorch

Tasks

Scene UnderstandingSurface Normal Estimation

Methods 이 논문이 사용한 방법론

Gravity Gravity is a kinematic approach to optimization based on gradients.

Similar Papers 제목 키워드 기반

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation

2026-07-16 · Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu, Xu Zheng 외 arxiv

Egocentric Visual Question Answering (VQA) has attracted widespread attention as an important task for enabling Multimodal Large Language Models (MLLMs) to interact with the real world. However, existing MLLMs struggle t…

Visual Question AnsweringSpatial Reasoning

Understanding Dynamic Scenes in Ego Centric 4D Point Clouds

2025-08-10 · Junsheng Huang, Shengyu Hao, Bocheng Hu, Hongwei Wang 외 arxiv

Understanding dynamic 4D scenes from an egocentric perspective-modeling changes in 3D spatial structure over time-is crucial for human-machine interaction, autonomous navigation, and embodied intelligence. While existing…

Trajectory PredictionScene UnderstandingPoint Clouds

EgoSplat: Open-Vocabulary Egocentric Scene Understanding with Language Embedded 3D Gaussian Splatting

2025-03-14 · Di Li, Jie Feng, Jiahao Chen, Weisheng Dong 외

Egocentric scenes exhibit frequent occlusions, varied viewpoints, and dynamic interactions compared to typical scene understanding tasks. Occlusions and varied viewpoints can lead to multi-view semantic inconsistencies, …

Scene UnderstandingSegmentation

Aria-NeRF: Multimodal Egocentric View Synthesis

2023-11-11 · Jiankai Sun, Jianing Qiu, Chuanyang Zheng, John Tucker 외

We seek to accelerate research in developing rich, multimodal scene models trained from egocentric data, based on differentiable volumetric ray-tracing inspired by Neural Radiance Fields (NeRFs). The construction of a Ne…

NeRF

Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks

2025-10-29 · Xu Zheng, Zihao Dongfang, Lutao Jiang, Boyuan Zheng 외 arxiv

Humans possess spatial reasoning abilities that enable them to understand spaces through multimodal observations, such as vision and sound. Large multimodal reasoning models extend these abilities by learning to perceive…

Vision-Language NavigationVisual Question AnsweringMultimodal ReasoningSpatial Reasoning