paper-with-me

홈 › Papers

LEGO: Learning Edge with Geometry all at Once by Watching Videos

2018-03-15 · CVPR 2018 6 · Zhenheng Yang, Peng Wang, Yang Wang, Wei Xu, Ram Nevatia

Learning to estimate 3D geometry in a single image by watching unlabeled videos via deep convolutional network is attracting significant attention. In this paper, we introduce a "3D as-smooth-as-possible (3D-ASAP)" prior inside the pipeline, which enables joint estimation of edges and 3D scene, yielding results with significant improvement in accuracy for fine detailed structures. Specifically, we define the 3D-ASAP prior by requiring that any two points recovered in 3D from an image should lie on an existing planar surface if no other cues provided. We design an unsupervised framework that Learns Edges and Geometry (depth, normal) all at Once (LEGO). The predicted edges are embedded into depth and surface normal smoothness terms, where pixels without edges in-between are constrained to satisfy the prior. In our framework, the predicted depths, normals and edges are forced to be consistent all the time. We conduct experiments on KITTI to evaluate our estimated geometry and CityScapes to perform edge evaluation. We show that in all of the tasks, i.e.depth, normal and edge, our algorithm vastly outperforms other state-of-the-art (SOTA) algorithms, demonstrating the benefits of our approach.

📄 PDF Abstract BibTeX arXiv:1803.05648

Code (1)

zhenheny/LEGO tf

Tasks

3D geometryAll

Similar Papers 제목 키워드 기반

Unsupervised Learning of Geometry with Edge-aware Depth-Normal Consistency

2017-11-10 · Zhenheng Yang, Peng Wang, Wei Xu, Liang Zhao 외

Learning to reconstruct depths in a single image by watching unlabeled videos via deep convolutional network (DCN) is attracting significant attention in recent years. In this paper, we introduce a surface normal represe…

Depth Estimation

The Compositional Nature of Event Representations in the Human Brain

2015-05-25

How does the human brain represent simple compositions of constituents: actors, verbs, objects, directions, and locations? Subjects viewed videos during neuroimaging (fMRI) sessions from which sentential descriptions of …

Classification

3D Human Pose Perception from Egocentric Stereo Videos

2023-12-30 · CVPR 2024 1 · Hiroyasu Akada, Jian Wang, Vladislav Golyanik, Christian Theobalt

While head-mounted devices are becoming more compact, they provide egocentric views with significant self-occlusions of the device user. Hence, existing methods often fail to accurately estimate complex 3D poses from ego…

3D Human Pose Estimation3D Scene ReconstructionEgocentric Pose EstimationPose Estimation

Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models

2023-11-23 · Saman Motamed, Danda Pani Paudel, Luc van Gool

Text-to-Image (T2I) models excel at synthesizing concepts such as nouns, appearances, and styles. To enable customized content creation based on a few example images of a concept, methods such as Textual Inversion and Dr…

Language ModellingLarge Language ModelQuestion AnsweringVisual Question Answering+1

SWaT: Statistical Modeling of Video Watch Time through User Behavior Analysis

2024-08-14 · Shentao Yang, Haichuan Yang, Linna Du, Adithya Ganesh 외

The significance of estimating video watch time has been highlighted by the rising importance of (short) video recommendation, which has become a core product of mainstream social media platforms. Modeling video watch ti…

Binary Classification