paper-with-me

홈 › Papers

Cross-modal Active Complementary Learning with Self-refining Correspondence

2023-10-26 · NeurIPS 2023 11 · Yang Qin, Yuan Sun, Dezhong Peng, Joey Tianyi Zhou, Xi Peng, Peng Hu

Recently, image-text matching has attracted more and more attention from academia and industry, which is fundamental to understanding the latent correspondence across visual and textual modalities. However, most existing methods implicitly assume the training pairs are well-aligned while ignoring the ubiquitous annotation noise, a.k.a noisy correspondence (NC), thereby inevitably leading to a performance drop. Although some methods attempt to address such noise, they still face two challenging problems: excessive memorizing/overfitting and unreliable correction for NC, especially under high noise. To address the two problems, we propose a generalized Cross-modal Robust Complementary Learning framework (CRCL), which benefits from a novel Active Complementary Loss (ACL) and an efficient Self-refining Correspondence Correction (SCC) to improve the robustness of existing methods. Specifically, ACL exploits active and complementary learning losses to reduce the risk of providing erroneous supervision, leading to theoretically and experimentally demonstrated robustness against NC. SCC utilizes multiple self-refining processes with momentum correction to enlarge the receptive field for correcting correspondences, thereby alleviating error accumulation and achieving accurate and stable corrections. We carry out extensive experiments on three image-text benchmarks, i.e., Flickr30K, MS-COCO, and CC152K, to verify the superior robustness of our CRCL against synthetic and real-world noisy correspondences.

📄 PDF Abstract BibTeX arXiv:2310.17468

Code (1)

qinyang79/crcl 공식 구현 pytorch

Tasks

Cross-modal retrieval with noisy correspondenceImage-text matchingText Matching

Similar Papers 제목 키워드 기반

Communication Policy Evolution for Proactive LLM Agents

2026-06-12 · Xinbei Ma, Jiyang Qiu, Yao Yao, Zheng Wu 외 arxiv

LLM agents have rapidly evolved into autonomous systems, yet a persistent information gap remains between users and agents: communication is costly, while users' identical preferences further limit information exchange. …

Image Understands Point Cloud: Weakly Supervised 3D Semantic Segmentation via Association Learning

2022-09-16 · Tianfang Sun, Zhizhong Zhang, Xin Tan, Yanyun Qu 외

Weakly supervised point cloud semantic segmentation methods that require 1\% or fewer labels, hoping to realize almost the same performance as fully supervised approaches, which recently, have attracted extensive researc…

3D Semantic SegmentationPseudo LabelSemantic SegmentationSuperpixels+1

Virtual Multi-Modality Self-Supervised Foreground Matting for Human-Object Interaction

2021-10-07 · ICCV 2021 10 · Bo Xu, Han Huang, Cheng Lu, Ziwen Li 외

Most existing human matting algorithms tried to separate pure human-only foreground from the background. In this paper, we propose a Virtual Multi-modality Foreground Matting (VMFM) method to learn human-object interacti…

DecoderHuman-Object Interaction DetectionImage Matting

CrossTracker: Robust Multi-modal 3D Multi-Object Tracking via Cross Correction

2024-11-28 · Lipeng Gu, Xuefeng Yan, Weiming Wang, Honghua Chen 외

The fusion of camera- and LiDAR-based detections offers a promising solution to mitigate tracking failures in 3D multi-object tracking (MOT). However, existing methods predominantly exploit camera detections to correct t…

3D Multi-Object TrackingMulti-Object TrackingObject Tracking

Shared Cross-Modal Trajectory Prediction for Autonomous Driving

2020-04-01 · Chiho Choi, Joon Hee Choi, Srikanth Malla, Jiachen Li

Predicting future trajectories of traffic agents in highly interactive environments is an essential and challenging problem for the safe operation of autonomous driving systems. On the basis of the fact that self-driving…

Autonomous DrivingFuture predictionPredictionTrajectory Prediction