paper-with-me

홈 › Papers

VidLBEval: Benchmarking and Mitigating Language Bias in Video-Involved LVLMs

2025-02-23 · Yiming Yang, Yangyang Guo, Hui Lu, Yan Wang

Recently, Large Vision-Language Models (LVLMs) have made significant strides across diverse multimodal tasks and benchmarks. This paper reveals a largely under-explored problem from existing video-involved LVLMs - language bias, where models tend to prioritize language over video and thus result in incorrect responses. To address this research gap, we first collect a Video Language Bias Evaluation Benchmark, which is specifically designed to assess the language bias in video-involved LVLMs through two key tasks: ambiguous video contrast and interrogative question probing. Accordingly, we design accompanied evaluation metrics that aim to penalize LVLMs being biased by language. In addition, we also propose Multi-branch Contrastive Decoding (MCD), introducing two expert branches to simultaneously counteract language bias potentially generated by the amateur text-only branch. Our experiments demonstrate that i) existing video-involved LVLMs, including both proprietary and open-sourced, are largely limited by the language bias problem; ii) our MCD can effectively mitigate this issue and maintain general-purpose capabilities in various video-involved LVLMs without any additional retraining or alteration to model architectures.

📄 PDF Abstract BibTeX arXiv:2502.16602

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

Frame Sampling Strategies Matter: A Benchmark for small vision language models

2025-09-18 · Marija Brkic, Anas Filali Razzouki, Yannis Tevissen, Khalil Guetari 외 arxiv

Comparing vision language models on videos is particularly complex, as the performances is jointly determined by the model's visual representation capacity and the frame-sampling strategy used to construct the input. Cur…

EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness

2026-05-01 · Yueru Sun, Yimeng Zhang, Haoyu Gu, Nuo Chen 외 arxiv

Multimodal Emotion Recognition (MER) is critical for interpreting real-world interactions. While Multimodal Large Language Models (MLLM) have shown promise in MER, their internal decision-making mechanisms under modality…

Multimodal Emotion Recognition

Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning

2025-01-09 · CVPR 2025 1 · Huabin Liu, Filip Ilievski, Cees G. M. Snoek

This paper proposes the first video-grounded entailment tree reasoning method for commonsense video question answering (VQA). Despite the remarkable progress of large visual-language models (VLMs), there are growing conc…

BenchmarkingQuestion AnsweringVideo Question AnsweringVisual Question Answering (VQA)

debiaSAE: Benchmarking and Mitigating Vision-Language Model Bias

2024-10-17 · Kuleen Sasse, Shan Chen, Jackson Pond, Danielle Bitterman 외

As Vision Language Models (VLMs) gain widespread use, their fairness remains under-explored. In this paper, we analyze demographic biases across five models and six datasets. We find that portrait datasets like UTKFace a…

BenchmarkingBias DetectionFairnessLanguage Modeling+1

Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models

2024-10-18 · Olga Loginova, Oleksandr Bezrukov, Alexey Kravets

Evaluating Video Language Models (VLMs) is a challenging task. Due to its transparency, Multiple-Choice Question Answering (MCQA) is widely used to measure the performance of these models through accuracy. However, exist…

FairnessMultiple-choiceMultiple Choice Question Answering (MCQA)Question Answering+1