paper-with-me

홈 › Papers

Attentional Separation-and-Aggregation Network for Self-supervised Depth-Pose Learning in Dynamic Scenes

2020-11-18 · Feng Gao, Jincheng Yu, Hao Shen, Yu Wang, Huazhong Yang

Learning depth and ego-motion from unlabeled videos via self-supervision from epipolar projection can improve the robustness and accuracy of the 3D perception and localization of vision-based robots. However, the rigid projection computed by ego-motion cannot represent all scene points, such as points on moving objects, leading to false guidance in these regions. To address this problem, we propose an Attentional Separation-and-Aggregation Network (ASANet), which can learn to distinguish and extract the scene's static and dynamic characteristics via the attention mechanism. We further propose a novel MotionNet with an ASANet as the encoder, followed by two separate decoders, to estimate the camera's ego-motion and the scene's dynamic motion field. Then, we introduce an auto-selecting approach to detect the moving objects for dynamic-aware learning automatically. Empirical experiments demonstrate that our method can achieve the state-of-the-art performance on the KITTI benchmark.

📄 PDF Abstract BibTeX arXiv:2011.09369

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

MotionNet MotionNet is a system for joint perception and motion prediction based on a bird's eye view (BEV) map, which encodes the object category and motion information from 3D point…

Similar Papers 제목 키워드 기반

Self-distilled Feature Aggregation for Self-supervised Monocular Depth Estimation

2022-09-15 · Zhengming Zhou, Qiulei Dong

Self-supervised monocular depth estimation has received much attention recently in computer vision. Most of the existing works in literature aggregate multi-scale features for depth prediction via either straightforward …

Depth EstimationDepth PredictionMonocular Depth Estimation

BaseBoostDepth: Exploiting Larger Baselines For Self-supervised Monocular Depth Estimation

2024-07-29 · Kieran Saunders, Luis J. Manso, George Vogiatzis

In the domain of multi-baseline stereo, the conventional understanding is that, in general, increasing baseline separation substantially enhances the accuracy of depth estimation. However, prevailing self-supervised dept…

Depth EstimationMonocular Depth EstimationPose Estimation

Why Self-Attention? A Targeted Evaluation of Neural Machine Translation Architectures

2018-08-27 · EMNLP 2018 10 · Gongbo Tang, Mathias Müller, Annette Rios, Rico Sennrich

Recently, non-recurrent architectures (convolutional, self-attentional) have outperformed RNNs in neural machine translation. CNNs and self-attentional networks can connect distant words via shorter network paths than RN…

Machine TranslationTranslationWord Sense Disambiguation

Sharp Bounds for Poly-GNNs and the Effect of Graph Noise

2024-07-28 · Luciano Vinas, Arash A. Amini

We investigate the classification performance of graph neural networks with graph-polynomial features, poly-GNNs, on the problem of semi-supervised node classification. We analyze poly-GNNs under a general contextual sto…

Node ClassificationStochastic Block Model

Auditory Separation of a Conversation from Background via Attentional Gating

2019-05-26 · Shariq Mobin, Bruno Olshausen

We present a model for separating a set of voices out of a sound mixture containing an unknown number of sources. Our Attentional Gating Network (AGN) uses a variable attentional context to specify which speakers in the …

Speaker Separation