paper-with-me

홈 › Papers

Exploiting Pseudo Image Captions for Multimodal Summarization

2023-05-09 · Chaoya Jiang, Rui Xie, Wei Ye, Jinan Sun, Shikun Zhang

Cross-modal contrastive learning in vision language pretraining (VLP) faces the challenge of (partial) false negatives. In this paper, we study this problem from the perspective of Mutual Information (MI) optimization. It is common sense that InfoNCE loss used in contrastive learning will maximize the lower bound of MI between anchors and their positives, while we theoretically prove that MI involving negatives also matters when noises commonly exist. Guided by a more general lower bound form for optimization, we propose a contrastive learning strategy regulated by progressively refined cross-modal similarity, to more accurately optimize MI between an image/text anchor and its negative texts/images instead of improperly minimizing it. Our method performs competitively on four downstream cross-modal tasks and systematically balances the beneficial and harmful effects of (partial) false negative samples under theoretical guidance.

📄 PDF Abstract BibTeX arXiv:2305.05496

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense ReasoningContrastive LearningImage Captioning

Methods 이 논문이 사용한 방법론

InfoNCE 설명 없음
Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

MMCIG: Multimodal Cover Image Generation for Text-only Documents and Its Dataset Construction via Pseudo-labeling

2025-08-24 · Hyeyeon Kim, Sungwoo Han, Jingun Kwon, Hidetaka Kamigaito 외 arxiv

In this study, we introduce a novel cover image generation task that produces both a concise summary and a visually corresponding image from a given text-only document. Because no existing datasets are available for this…

Image Generation

UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation

2021-09-13 · Zhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang 외

With the rapid increase of multimedia data, a large body of literature has emerged to work on multimodal summarization, the majority of which target at refining salient information from textual and visual modalities to o…

Abstractive Text SummarizationDecoderImage CaptioningKnowledge Distillation+1

SD-RSIC: Summarization Driven Deep Remote Sensing Image Captioning

2020-06-15 · Gencer Sumbul, Sonali Nayak, Begüm Demir

Deep neural networks (DNNs) have been recently found popular for image captioning problems in remote sensing (RS). Existing DNN based approaches rely on the availability of a training set made up of a high number of RS i…

Image Captioning

BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models

2025-10-23 · Ziheng Zhang, Xinyue Ma, Arpita Chowdhury, Elizabeth G. Campolongo 외 arxiv

This work investigates descriptive captions as an additional source of supervision for biological multimodal foundation models. Images and captions can be viewed as complementary samples from the latent morphospace of a …

Image Retrieval

MM-AVS: A Full-Scale Dataset for Multi-modal Summarization

2021-06-01 · NAACL 2021 4 · Xiyan Fu, Jun Wang, Zhenglu Yang

Multimodal summarization becomes increasingly significant as it is the basis for question answering, Web search, and many other downstream tasks. However, its learning materials have been lacking a holistic organization …

Question Answering