paper-with-me

Papers

Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification

2025-11-22 · Yangyang Liu, Yuhao Wang, Pingping Zhang arxiv

Multi-modal object Re-IDentification (ReID) is devoted to retrieving specific objects through the exploitation of complementary multi-modal image information. Existing methods mainly concentrate on the fusion of multi-modal features, yet neglecting the background interference. Besides, current multi-modal fusion methods often focus on aligning modality pairs but suffer from multi-modal consistency alignment. To address these issues, we propose a novel selective interaction and global-local alignment framework called Signal for multi-modal object ReID. Specifically, we first propose a Selective Interaction Module (SIM) to select important patch tokens with intra-modal and inter-modal information. These important patch tokens engage in the interaction with class tokens, thereby yielding more discriminative features. Then, we propose a Global Alignment Module (GAM) to simultaneously align multi-modal features by minimizing the volume of 3D polyhedra in the gramian space. Meanwhile, we propose a Local Alignment Module (LAM) to align local features in a shift-aware manner. With these modules, our proposed framework could extract more discriminative features for object ReID. Extensive experiments on three multi-modal object ReID benchmarks (i.e., RGBNT201, RGBNT100, MSVR310) validate the effectiveness of our method. The source code is available at https://github.com/010129/Signal.

📄 PDF Abstract BibTeX arXiv:2511.17965

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sound Source Localization for Human-Robot Interaction in Outdoor Environments

2025-07-29 · Victor Liu, Timothy Du, Jordy Sehn, Jack Collier 외 arxiv

This paper presents a sound source localization strategy that relies on a microphone array embedded in an unmanned ground vehicle and an asynchronous close-talking microphone near the operator. A signal coarse alignment …

Sound Source Localization

Guiding Federated Graph Recommendation with LLM-encoded knowledge

2026-06-13 · Thi Minh Chau Nguyen, Hien Trang Nguyen, Duc Anh Nguyen, Van Ho-Long 외 arxiv

Graph-based recommender systems are highly effective at extracting collaborative signals from user--item interactions, and federated learning (FL) allows these models to be trained while preserving user privacy. However,…

Federated Learning

SGANet: Semantic and Geometric Alignment for Multimodal Multi-view Anomaly Detection

2026-04-07 · Letian Bai, Chengyu Tao, Juan Du arxiv

Multi-view anomaly detection aims to identify surface defects on complex objects using observations captured from multiple viewpoints. However, existing unsupervised methods often suffer from feature inconsistency arisin…

Anomaly Detection

InterFormer: Interactive Local and Global Features Fusion for Automatic Speech Recognition

2023-05-24 · Zhi-Hao Lai, Tian-Hao Zhang, Qi Liu, Xinyuan Qian 외

The local and global features are both essential for automatic speech recognition (ASR). Many recent methods have verified that simply combining local and global features can further promote ASR performance. However, the…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Context-Aware Attention Network for Image-Text Retrieval

2020-06-01 · CVPR 2020 6 · Qi Zhang, Zhen Lei, Zhaoxiang Zhang, Stan Z. Li

As a typical cross-modal problem, image-text bi-directional retrieval relies heavily on the joint embedding learning and similarity measure for each image-text pair. It remains challenging because prior works seldom expl…

Image-text RetrievalRetrievalText Retrieval