paper-with-me

홈 › Papers

Learning Spatial Common Sense with Geometry-Aware Recurrent Networks

2018-12-31 · CVPR 2019 6 · Hsiao-Yu Fish Tung, Ricson Cheng, Katerina Fragkiadaki

We integrate two powerful ideas, geometry and deep visual representation learning, into recurrent network architectures for mobile visual scene understanding. The proposed networks learn to "lift" and integrate 2D visual features over time into latent 3D feature maps of the scene. They are equipped with differentiable geometric operations, such as projection, unprojection, egomotion estimation and stabilization, in order to compute a geometrically-consistent mapping between the world scene and their 3D latent feature state. We train the proposed architectures to predict novel camera views given short frame sequences as input. Their predictions strongly generalize to scenes with a novel number of objects, appearances and configurations; they greatly outperform previous works that do not consider egomotion stabilization or a space-aware latent feature state. We train the proposed architectures to detect and segment objects in 3D using the latent 3D feature map as input--as opposed to per frame features. The resulting object detections persist over time: they continue to exist even when an object gets occluded or leaves the field of view. Our experiments suggest the proposed space-aware latent feature memory and egomotion-stabilized convolutions are essential architectural choices for spatial common sense to emerge in artificial embodied visual agents.

📄 PDF Abstract BibTeX arXiv:1901.00003

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense ReasoningRepresentation LearningScene Understanding

Similar Papers 제목 키워드 기반

Paragraph-level Commonsense Transformers with Recurrent Memory

2020-10-04 · Saadia Gabriel, Chandra Bhagavatula, Vered Shwartz, Ronan Le Bras 외

Human understanding of narrative texts requires making commonsense inferences beyond what is stated explicitly in the text. A recent model, COMET, can generate such implicit commonsense inferences along several dimension…

SentenceWorld Knowledge

Global-Aware Enhanced Spatial-Temporal Graph Recurrent Networks: A New Framework For Traffic Flow Prediction

2024-01-07 · Haiyang Liu, Chunjiang Zhu, Detian Zhang

Traffic flow prediction plays a crucial role in alleviating traffic congestion and enhancing transport efficiency. While combining graph convolution networks with recurrent neural networks for spatial-temporal modeling i…

Graph Neural NetworkPredictionTraffic Prediction

LA-Sign: Looped Transformers with Geometry-aware Alignment for Skeleton-based Sign Language Recognition

2026-03-30 · Muxin Pu, Mei Kuan Lim, Chun Yong Chong, Chen Change Loy arxiv

Skeleton-based isolated sign language recognition (ISLR) demands fine-grained understanding of articulated motion across multiple spatial scales, from subtle finger movements to global body dynamics. Existing approaches …

Sign Language RecognitionRepresentation Learning

Things not Written in Text: Exploring Spatial Commonsense from Visual Signals

2022-03-15 · ACL 2022 5 · Xiao Liu, Da Yin, Yansong Feng, Dongyan Zhao

Spatial commonsense, the knowledge about spatial position and relationship between objects (like the relative size of a lion and a girl, and the position of a boy relative to a bicycle when cycling), is an important part…

Image GenerationNatural Language UnderstandingPosition

Things not Written in Text: Exploring Spatial Commonsense from Visual Signals

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Spatial commonsense, the knowledge about spatial position and relationship between objects (like the relative size of a lion and a girl, and the position of a boy relative to a bicycle when cycling), is an important part…

Image GenerationNatural Language UnderstandingPosition