paper-with-me

홈 › Papers

Disentangled Noisy Correspondence Learning

2024-08-10 · Zhuohang Dang, Minnan Luo, Jihong Wang, Chengyou Jia, Haochen Han, Herun Wan, Guang Dai, Xiaojun Chang, Jingdong Wang

Cross-modal retrieval is crucial in understanding latent correspondences across modalities. However, existing methods implicitly assume well-matched training data, which is impractical as real-world data inevitably involves imperfect alignments, i.e., noisy correspondences. Although some works explore similarity-based strategies to address such noise, they suffer from sub-optimal similarity predictions influenced by modality-exclusive information (MEI), e.g., background noise in images and abstract definitions in texts. This issue arises as MEI is not shared across modalities, thus aligning it in training can markedly mislead similarity predictions. Moreover, although intuitive, directly applying previous cross-modal disentanglement methods suffers from limited noise tolerance and disentanglement efficacy. Inspired by the robustness of information bottlenecks against noise, we introduce DisNCL, a novel information-theoretic framework for feature Disentanglement in Noisy Correspondence Learning, to adaptively balance the extraction of MII and MEI with certifiable optimal cross-modal disentanglement efficacy. DisNCL then enhances similarity predictions in modality-invariant subspace, thereby greatly boosting similarity-based alleviation strategy for noisy correspondences. Furthermore, DisNCL introduces soft matching targets to model noisy many-to-many relationships inherent in multi-modal input for noise-robust and accurate cross-modal alignment. Extensive experiments confirm DisNCL's efficacy by 2% average recall improvement. Mutual information estimation and visualization results show that DisNCL learns meaningful MII/MEI subspaces, validating our theoretical analyses.

📄 PDF Abstract BibTeX arXiv:2408.05503

Code (0)

등록된 구현이 없습니다.

Tasks

cross-modal alignmentCross-Modal RetrievalDisentanglementMutual Information Estimation

Methods 이 논문이 사용한 방법론

MEI MEI introduces the *multi-partition embedding interaction* technique with block term tensor format to systematically address the efficiency--expressiveness trade-off in…

Similar Papers 제목 키워드 기반

Breaking Through the Noisy Correspondence: A Robust Model for Image-Text Matching

2024-04-29 · ACM Transactions on Information Systems 2024 4 · Haitao Shi, Meng Liu, Xiaoxuan Mu, Xuemeng Song 외

Unleashing the power of image-text matching in real-world applications is hampered by noisy correspondence. Manually curating high-quality datasets is expensive and time-consuming, and datasets generated using diffusion …

Cross-modal retrieval with noisy correspondenceImage-text matchingText Matching

Curriculum Disentangled Recommendation with Noisy Multi-feedback

2021-12-01 · NeurIPS 2021 12 · Hong Chen, Yudong Chen, Xin Wang, Ruobing Xie 외

Learning disentangled representations for user intentions from multi-feedback (i.e., positive and negative feedback) can enhance the accuracy and explainability of recommendation algorithms. However, learning such disent…

DenoisingRepresentation Learning

Learning with Noisy Correspondence for Cross-modal Matching

2021-12-01 · NeurIPS 2021 12 · Zhenyu Huang, guocheng niu, Xiao Liu, Wenbiao Ding 외

Cross-modal matching, which aims to establish the correspondence between two different modalities, is fundamental to a variety of tasks such as cross-modal retrieval and vision-and-language understanding. Although a huge…

Cross-Modal RetrievalCross-modal retrieval with noisy correspondenceImage-text matchingMemorization+2

Mitigating Noisy Correspondence by Geometrical Structure Consistency Learning

2024-05-27 · CVPR 2024 1 · Zihua Zhao, Mengxi Chen, Tianjie Dai, Jiangchao Yao 외

Noisy correspondence that refers to mismatches in cross-modal data pairs, is prevalent on human-annotated or web-crawled datasets. Prior approaches to leverage such data mainly consider the application of uni-modal noisy…

Cross-modal retrieval with noisy correspondence

Learning From Noisy Correspondence With Tri-Partition for Cross-Modal Matching

2023-09-22 · IEEE Transactions on Multimedia 2023 9 · Zerun Feng, Zhimin Zeng, Caili Guo, Zheng Li 외

Due to high labeling cost, it is inevitable to introduce a certain proportion of noisy correspondence into visual-text datasets, resulting in poor model robustness for cross-modal matching. Although recent methods divide…

Cross-modal retrieval with noisy correspondenceMemorizationSemantic correspondenceText Matching+1