paper-with-me

홈 › Papers

Fusing Structure from Motion and Simulation-Augmented Pose Regression from Optical Flow for Challenging Indoor Environments

2023-04-14 · Felix Ott, Lucas Heublein, David Rügamer, Bernd Bischl, Christopher Mutschler

The localization of objects is a crucial task in various applications such as robotics, virtual and augmented reality, and the transportation of goods in warehouses. Recent advances in deep learning have enabled the localization using monocular visual cameras. While structure from motion (SfM) predicts the absolute pose from a point cloud, absolute pose regression (APR) methods learn a semantic understanding of the environment through neural networks. However, both fields face challenges caused by the environment such as motion blur, lighting changes, repetitive patterns, and feature-less structures. This study aims to address these challenges by incorporating additional information and regularizing the absolute pose using relative pose regression (RPR) methods. RPR methods suffer under different challenges, i.e., motion blur. The optical flow between consecutive images is computed using the Lucas-Kanade algorithm, and the relative pose is predicted using an auxiliary small recurrent convolutional network. The fusion of absolute and relative poses is a complex task due to the mismatch between the global and local coordinate systems. State-of-the-art methods fusing absolute and relative poses use pose graph optimization (PGO) to regularize the absolute pose predictions using relative poses. In this work, we propose recurrent fusion networks to optimally align absolute and relative pose predictions to improve the absolute pose prediction. We evaluate eight different recurrent units and construct a simulation environment to pre-train the APR and RPR networks for better generalized training. Additionally, we record a large database of different scenarios in a challenging large-scale indoor environment that mimics a warehouse with transportation robots. We conduct hyperparameter searches and experiments to show the effectiveness of our recurrent fusion method compared to PGO.

📄 PDF Abstract BibTeX arXiv:2304.07250

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Flow EstimationPose Predictionregression

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Fusing Convolutional Neural Network and Geometric Constraint for Image-based Indoor Localization

2022-01-05 · Jingwei Song, Mitesh Patel, Maani Ghaffari

This paper proposes a new image-based localization framework that explicitly localizes the camera/robot by fusing Convolutional Neural Network (CNN) and sequential images' geometric constraints. The camera is localized u…

Image-Based LocalizationIndoor LocalizationPose Prediction

CMATH: Cross-Modality Augmented Transformer with Hierarchical Variational Distillation for Multimodal Emotion Recognition in Conversation

2024-11-15 · Xiaofei Zhu, Jiawei Cheng, Zhou Yang, Zhuo Chen 외

Multimodal emotion recognition in conversation (MER) aims to accurately identify emotions in conversational utterances by integrating multimodal information. Previous methods usually treat multimodal information as equal…

Emotion RecognitionEmotion Recognition in ConversationMultimodal Emotion Recognitionmultimodal interaction

Disparity-Augmented Trajectories for Human Activity Recognition

2019-05-14 · Pejman Habashi, Boubakeur Boufama, Imran Shafiq Ahmad

Numerous methods for human activity recognition have been proposed in the past two decades. Many of these methods are based on sparse representation, which describes the whole video content by a set of local features. Tr…

3D ReconstructionActivity RecognitionHuman Activity Recognition

Augmented LRFS-based Filter: Holistic Tracking of Group Objects

2024-03-20 · Chaoqun Yang, Xiaowei Liang, Zhiguo Shi, Heng Zhang 외

This paper addresses the problem of group target tracking (GTT), wherein multiple closely spaced targets within a group pose a coordinated motion. To improve the tracking performance, the labeled random finite sets (LRFS…

GLDiTalker: Speech-Driven 3D Facial Animation with Graph Latent Diffusion Transformer

2024-08-03 · Yihong Lin, Zhaoxin Fan, Xianjia Wu, Lingyu Xiong 외

Speech-driven talking head generation is a critical yet challenging task with applications in augmented reality and virtual human modeling. While recent approaches using autoregressive and diffusion-based models have ach…

DiversityTalking Head Generation