paper-with-me

홈 › Papers

AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering

2025-08-25 · Kang Zeng, Guojin Zhong, Jintao Cheng, Jin Yuan, Zhiyong Li arxiv

The advancement of Multimodal Large Language Models (MLLMs) has driven significant progress in Visual Question Answering (VQA), evolving from Single to Multi Image VQA (MVQA). However, the increased number of images in MVQA inevitably introduces substantial visual redundancy that is irrelevant to question answering, negatively impacting both accuracy and efficiency. To address this issue, existing methods lack flexibility in controlling the number of compressed visual tokens and tend to produce discrete visual fragments, which hinder MLLMs' ability to comprehend images holistically. In this paper, we propose a straightforward yet universal Adaptive Visual Anchoring strategy, which can be seamlessly integrated into existing MLLMs, offering significant accuracy improvements through adaptive compression. Meanwhile, to balance the results derived from both global and compressed visual input, we further introduce a novel collaborative decoding mechanism, enabling optimal performance. Extensive experiments validate the effectiveness of our method, demonstrating consistent performance improvements across various MLLMs. The code will be publicly available.

📄 PDF Abstract BibTeX arXiv:2508.17860

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

Discrete Auto-regressive Variational Attention Models for Text Modeling

2021-06-16 · Xianghong Fang, Haoli Bai, Jian Li, Zenglin Xu 외

Variational autoencoders (VAEs) have been widely applied for text modeling. In practice, however, they are troubled by two challenges: information underrepresentation and posterior collapse. The former arises as only the…

Language ModelingLanguage Modelling

GaVaMoE: Gaussian-Variational Gated Mixture of Experts for Explainable Recommendation

2024-10-15 · Fei Tang, Yongliang Shen, Hang Zhang, Zeqi Tan 외

Large language model-based explainable recommendation (LLM-based ER) systems show promise in generating human-like explanations for recommendations. However, they face challenges in modeling user-item collaborative prefe…

Explainable RecommendationLanguage ModellingLarge Language ModelMixture-of-Experts

VaViM and VaVAM: Autonomous Driving through Video Generative Modeling

2025-02-21 · Florent Bartoccioni, Elias Ramzi, Victor Besnier, Shashanka Venkataramanan 외

We explore the potential of large-scale generative video models for autonomous driving, introducing an open-source auto-regressive video model (VaViM) and its companion video-action model (VaVAM) to investigate how video…

Autonomous DrivingImitation Learning

An Audio-Visual Attention Based Multimodal Network for Fake Talking Face Videos Detection

2022-03-10 · Ganglai Wang, Peng Zhang, Lei Xie, Wei Huang 외

DeepFake based digital facial forgery is threatening the public media security, especially when lip manipulation has been used in talking face generation, the difficulty of fake video detection is further improved. By on…

Decision MakingFace DetectionFace GenerationFace Swapping+1

FTFDNet: Learning to Detect Talking Face Video Manipulation with Tri-Modality Interaction

2023-07-08 · Ganglai Wang, Peng Zhang, Junwen Xiong, Feihan Yang 외

DeepFake based digital facial forgery is threatening public media security, especially when lip manipulation has been used in talking face generation, and the difficulty of fake video detection is further improved. By on…

Face DetectionFace GenerationFace SwappingOptical Flow Estimation+1