paper-with-me

Papers

MIR-GAN: Refining Frame-Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech Recognition

2023-06-18 · Yuchen Hu, Chen Chen, Ruizhe Li, Heqing Zou, Eng Siong Chng

Audio-visual speech recognition (AVSR) attracts a surge of research interest recently by leveraging multimodal signals to understand human speech. Mainstream approaches addressing this task have developed sophisticated architectures and techniques for multi-modality fusion and representation learning. However, the natural heterogeneity of different modalities causes distribution gap between their representations, making it challenging to fuse them. In this paper, we aim to learn the shared representations across modalities to bridge their gap. Different from existing similar methods on other multimodal tasks like sentiment analysis, we focus on the temporal contextual dependencies considering the sequence-to-sequence task setting of AVSR. In particular, we propose an adversarial network to refine frame-level modality-invariant representations (MIR-GAN), which captures the commonality across modalities to ease the subsequent multimodal fusion process. Extensive experiments on public benchmarks LRS3 and LRS2 show that our approach outperforms the state-of-the-arts.

📄 PDF Abstract BibTeX arXiv:2306.10567

Code (1)

yuchen005/mir-gan 공식 구현 pytorch

Tasks

Audio-Visual Speech RecognitionRepresentation LearningSentiment Analysisspeech-recognitionSpeech RecognitionVisual Speech Recognition

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Semantic-Guided Multimodal Sentiment Decoding with Adversarial Temporal-Invariant Learning

2024-08-30 · Guoyang Xu, Junqi Xue, Yuxin Liu, ZiRui Wang 외

Multimodal sentiment analysis aims to learn representations from different modalities to identify human emotions. However, existing works often neglect the frame-level redundancy inherent in continuous time series, resul…

Multimodal Sentiment AnalysisSentiment AnalysisTime Series

Learning Language-Driven Sequence-Level Modal-Invariant Representations for Video-Based Visible-Infrared Person Re-Identification

2026-01-17 · Xiaomei Yang, Antai Liu, Xizhan Gao, Fa Zhu 외 arxiv

The core of video-based visible-infrared person re-identification (VVI-ReID) lies in learning sequence-level modal-invariant representations across different modalities. Recent research tends to use modality-shared langu…

Person Re-IdentificationRepresentation Learning

Modality-Adaptive Mixup and Invariant Decomposition for RGB-Infrared Person Re-Identification

2022-03-03 · Zhipeng Huang, Jiawei Liu, Liang Li, Kecheng Zheng 외

RGB-infrared person re-identification is an emerging cross-modality re-identification task, which is very challenging due to significant modality discrepancy between RGB and infrared images. In this work, we propose a no…

Deep Reinforcement LearningPerson Re-Identification

Dynamic Modality-Camera Invariant Clustering for Unsupervised Visible-Infrared Person Re-identification

2024-12-11 · Yiming Yang, Weipeng Hu, Haifeng Hu

Unsupervised learning visible-infrared person re-identification (USL-VI-ReID) offers a more flexible and cost-effective alternative compared to supervised methods. This field has gained increasing attention due to its pr…

ClusteringContrastive LearningPerson Re-Identification

DeMIAN: Deep Modality Invariant Adversarial Network

2016-12-23 · Kuniaki Saito, Yusuke Mukuta, Yoshitaka Ushiku, Tatsuya Harada

Obtaining common representations from different modalities is important in that they are interchangeable with each other in a classification problem. For example, we can train a classifier on image features in the common…

Domain AdaptationGeneral ClassificationRepresentation LearningZero-Shot Learning