paper-with-me

홈 › Papers

VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning

2026-03-01 · Mingkang Dong, Hongyi Cai, Jie Li, Sifan Zhou, Bin Ren, Kunyu Peng, Yuqian Fu arxiv

The effectiveness of multimodal instruction tuning depends not only on dataset scale, but critically on whether training samples genuinely require visual reasoning. However, existing instruction datasets often contain a substantial portion of visually redundant samples (solvable from text alone), as well as multimodally misaligned supervision that can degrade learning. To address this, we propose VisNec (Visual Necessity Score), a principled data selection framework that measures the marginal contribution of visual input during instruction tuning. By comparing predictive loss with and without visual context, VisNec identifies whether a training instance is vision-critical, redundant, or misaligned. To preserve task diversity, we combine VisNec with semantic clustering and select high-necessity samples within each cluster. Across 10 downstream benchmarks, training on only 15% of the LLaVA-665K dataset selected by VisNec achieves 100.2% of full-data performance. On the smaller Vision-Flan-186K dataset, our selection not only further reduces data size but also surpasses full-data training by 15.8%. These results demonstrate that measuring and leveraging visual necessity provides an effective solution for both efficient and robust multimodal instruction tuning. Codes and selected subsets will be released upon acceptance.

📄 PDF Abstract BibTeX arXiv:2603.01195

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

Exploring the Necessity of Visual Modality in Multimodal Machine Translation using Authentic Datasets

2024-04-09 · Zi Long, Zhenhao Tang, Xianghua Fu, Jian Chen 외

Recent research in the field of multimodal machine translation (MMT) has indicated that the visual modality is either dispensable or offers only marginal advantages. However, most of these conclusions are drawn from the …

Machine TranslationMultimodal Machine TranslationSentenceTranslation

Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos

2026-01-11 · Haodong Chen, Qiang Huang, Jiaqi Zhao, Qiuping Jiang 외 arxiv

Vision-Language Models (VLMs) are increasingly deployed in socially consequential settings, raising concerns about social bias driven by demographic cues. A central challenge in measuring such social bias is attribution …

Leveraging Entity Information for Cross-Modality Correlation Learning: The Entity-Guided Multimodal Summarization

2024-08-06 · Yanghai Zhang, Ye Liu, Shiwei Wu, Kai Zhang 외

The rapid increase in multimedia data has spurred advancements in Multimodal Summarization with Multimodal Output (MSMO), which aims to produce a multimodal summary that integrates both text and relevant images. The inhe…

Knowledge DistillationLanguage ModelingLanguage Modelling

What do Models Learn From Training on More Than Text? Measuring Visual Commonsense Knowledge

2022-05-14 · ACL 2022 5 · Lovisa Hagström, Richard Johansson

There are limitations in learning language from text alone. Therefore, recent focus has been on developing multimodal models. However, few benchmarks exist that can measure what language models learn about language from …

Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity

2026-04-06 · Jaeyoon Jung, Yejun Yoon, Kunwoo Park arxiv

Automated fact-checking is a crucial task that supports a responsible information ecosystem. While recent research has progressed from text-only to multimodal fact-checking, a prevailing assumption is that incorporating …