paper-with-me

홈 › Papers

Learning to Exploit Stability for 3D Scene Parsing

2018-12-01 · NeurIPS 2018 12 · Yilun Du, Zhijian Liu, Hector Basevi, Ales Leonardis, Bill Freeman, Josh Tenenbaum, Jiajun Wu

Human scene understanding uses a variety of visual and non-visual cues to perform inference on object types, poses, and relations. Physics is a rich and universal cue which we exploit to enhance scene understanding. We integrate the physical cue of stability into the learning process using a REINFORCE approach coupled to a physics engine, and apply this to the problem of producing the 3D bounding boxes and poses of objects in a scene. We first show that applying physics supervision to an existing scene understanding model increases performance, produces more stable predictions, and allows training to an equivalent performance level with fewer annotated training examples. We then present a novel architecture for 3D scene parsing named Prim R-CNN, learning to predict bounding boxes as well as their 3D size, translation, and rotation. With physics supervision, Prim R-CNN outperforms existing scene understanding approaches on this problem. Finally, we show that applying physics supervision on unlabeled real images improves real domain transfer of models training on synthetic data.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Scene ParsingScene UnderstandingTranslation

Methods 이 논문이 사용한 방법론

REINFORCE REINFORCE is a Monte Carlo variant of a policy gradient algorithm in reinforcement learning. The agent collects samples of an episode using its current policy, and uses it to…

Similar Papers 제목 키워드 기반

Pyramid Scene Parsing Network

2016-12-04 · CVPR 2017 7 · Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang 외

Scene parsing is challenging for unrestricted open vocabulary and diverse scenes. In this paper, we exploit the capability of global context information by different-region-based context aggregation through our pyramid p…

Dichotomous Image SegmentationImage ClassificationLesion SegmentationReal-Time Semantic Segmentation+4

Fully Exploiting Vision Foundation Model's Profound Prior Knowledge for Generalizable RGB-Depth Driving Scene Parsing

2025-02-10 · Sicen Guo, Tianyou Wen, Chuang-Wei Liu, Qijun Chen 외

Recent vision foundation models (VFMs), typically based on Vision Transformer (ViT), have significantly advanced numerous computer vision tasks. Despite their success in tasks focused solely on RGB images, the potential …

Depth EstimationDepth PredictionScene Parsing

FoveaNet: Perspective-aware Urban Scene Parsing

2017-08-08 · ICCV 2017 10 · Xin Li, Zequn Jie, Wei Wang, Changsong Liu 외

Parsing urban scene images benefits many applications, especially self-driving. Most of the current solutions employ generic image parsing models that treat all scales and locations in the images equally and do not consi…

Scene Parsing

Predicting Scene Parsing and Motion Dynamics in the Future

2017-11-09 · NeurIPS 2017 12 · Xiaojie Jin, Huaxin Xiao, Xiaohui Shen, Jimei Yang 외

The ability of predicting the future is important for intelligent systems, e.g. autonomous vehicles and robots to plan early and make decisions accordingly. Future scene parsing and optical flow estimation are two key ta…

Autonomous Vehiclesmotion predictionOptical Flow EstimationScene Parsing

VSPW: A Large-scale Dataset for Video Scene Parsing in the Wild

2021-06-19 · CVPR 2021 1 · Jiaxu Miao, Yunchao Wei, Yu Wu, Chen Liang 외

In this paper, we present a new dataset with the target of advancing the scene parsing task from images to videos. Our dataset aims to perform Video Scene Parsing in the Wild (VSPW), which covers a wide range of real…

4kScene Parsing