paper-with-me

Papers

Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models

2026-05-03 · Yuriel Ryan, Hei Man Ip, Adriel Kuek, Paul Pu Liang, Roy Ka-Wei Lee arxiv

Current vision language models face hallucination and robustness issues against ambiguous or corrupted modalities. We hypothesize that these issues can be addressed by exploiting the shared information between modalities to compensate for the impaired one. To this end, we analyze multimodal interactions -- redundant (shared), unique (exclusive), and synergistic (emergent) task-relevant information provided by the modalities -- to determine their impacts on model reliability. Specifically, amplifying redundant interactions would increase this exploitable shared information to resolve these issues; yet, modern instruction datasets often eliminate redundancies to prioritize visual grounding. We bridge this gap through a self-captioning workflow featuring a \textsc{Multimodal Interaction Gate}: a mechanism to convert unique interactions into redundant interactions. Our findings suggest that increasing redundancy can reduce visual induced errors by 38.3\% and improve consistency by 16.8\%.

📄 PDF Abstract BibTeX arXiv:2605.08145

Code (0)

등록된 구현이 없습니다.

Tasks

Face HallucinationVisual Grounding

Similar Papers 제목 키워드 기반

Multimodal Transformer with Multi-View Visual Representation for Image Captioning

2019-05-20 · Jun Yu, Jing Li, Zhou Yu, Qingming Huang

Image captioning aims to automatically generate a natural language description of a given image, and most state-of-the-art models have adopted an encoder-decoder framework. The framework consists of a convolution neural …

DecoderImage CaptioningMachine TranslationMultimodal Reasoning

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis

2024-12-04 · Davide Bucciarelli, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi 외

The task of image captioning demands an algorithm to generate natural language descriptions of visual inputs. Recent advancements have seen a convergence between image captioning research and the development of Large Lan…

Image CaptioningImage DescriptionPrompt Learning

On Advances in Text Generation from Images Beyond Captioning: A Case Study in Self-Rationalization

2022-05-24 · Shruti Palaskar, Akshita Bhagia, Yonatan Bisk, Florian Metze 외

Combining the visual modality with pretrained language models has been surprisingly effective for simple descriptive tasks such as image captioning. More general text generation however remains elusive. We take a step ba…

DescriptiveImage CaptioningNatural Language InferenceQuestion Answering+4

EchoChange: A Diffusion Language Model with Dual Pass Remasking for Factual Remote Sensing Disaster Change Captioning

2026-08-03 · Dongwei Sun, Bowen Yao, Yujie Zhang, Pei Liu 외 arxiv

Bi-temporal remote-sensing disaster change captioning often needs to identify sparse and spatially localized changes across large pre- and post-event scenes and then translate them into coherent, factual descriptions. Ho…

Quantifying Societal Bias Amplification in Image Captioning

2022-03-29 · CVPR 2022 1 · Yusuke Hirota, Yuta Nakashima, Noa Garcia

We study societal bias amplification in image captioning. Image captioning models have been shown to perpetuate gender and racial biases, however, metrics to measure, quantify, and evaluate the societal bias in captions …

AttributeImage Captioning