paper-with-me

홈 › Papers

Multi-modal Visual Place Recognition in Dynamics-Invariant Perception Space

2021-05-17 · Lin Wu, Teng Wang, Changyin Sun

Visual place recognition is one of the essential and challenging problems in the fields of robotics. In this letter, we for the first time explore the use of multi-modal fusion of semantic and visual modalities in dynamics-invariant space to improve place recognition in dynamic environments. We achieve this by first designing a novel deep learning architecture to generate the static semantic segmentation and recover the static image directly from the corresponding dynamic image. We then innovatively leverage the spatial-pyramid-matching model to encode the static semantic segmentation into feature vectors. In parallel, the static image is encoded using the popular Bag-of-words model. On the basis of the above multi-modal features, we finally measure the similarity between the query image and target landmark by the joint similarity of their semantic and visual codes. Extensive experiments demonstrate the effectiveness and robustness of the proposed approach for place recognition in dynamic environments.

📄 PDF Abstract BibTeX arXiv:2105.07800

Code (1)

fiftywu/Multimodal-VPR 공식 구현 pytorch

Tasks

SegmentationSemantic SegmentationVisual Place Recognition

Similar Papers 제목 키워드 기반

MSSPlace: Multi-Sensor Place Recognition with Visual and Text Semantics

2024-07-22 · Alexander Melekhin, Dmitry Yudin, Ilia Petryashin, Vitaly Bezuglyj

Place recognition is a challenging task in computer vision, crucial for enabling autonomous vehicles and robots to navigate previously visited environments. While significant progress has been made in learnable multimoda…

Autonomous VehiclesNavigateSemantic Segmentation

End-to-end Audiovisual Speech Recognition

2018-02-18 · IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2018 9 · Stavros Petridis, Themos Stafylakis, Pingchuan Ma, Feipeng Cai 외

Several end-to-end deep learning approaches have been recently presented which extract either audio or visual features from the input images or audio signals and perform speech recognition. However, research on end-to-en…

Lipreadingspeech-recognitionSpeech Recognition

Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition

2024-07-09 · Mingfang Zhang, Yifei HUANG, Ruicong Liu, Yoichi Sato

Compared with visual signals, Inertial Measurement Units (IMUs) placed on human limbs can capture accurate motion signals while being robust to lighting variation and occlusion. While these characteristics are intuitivel…

Action Recognition

OpenMPR: Recognize Places Using Multimodal Data for People with Visual Impairments

2019-09-15 · Ruiqi Cheng, Kaiwei Wang, Jian Bai, Zhijie Xu

Place recognition plays a crucial role in navigational assistance, and is also a challenging issue of assistive technology. The place recognition is prone to erroneous localization owing to various changes between databa…

MMS-VPR: Multimodal Street-Level Visual Place Recognition Dataset and Benchmark

2025-05-18 · Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao 외

Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, lack multimodal diversity and underrepresent dense, mixed-use street-level spaces, especially in non-Western urban contexts.…

Multimodal ReasoningVisual Place Recognition