paper-with-me

홈 › Papers

CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models

2026-03-22 · Nan Zhou, Huiqun Wang, Yaoyan Zheng, Di Huang arxiv

Multimodal large language models (MLLMs) achieve remarkable progress in cross-modal perception and reasoning, yet a fundamental question remains unresolved: should the vision encoder be fine-tuned or frozen? Despite the success of models such as LLaVA and Qwen-VL, inconsistent design choices and heterogeneous training setups hinder a unified understanding of visual fine-tuning (VFT) in MLLMs. Through a configuration-aligned benchmark, we find that existing VFT methods fail to consistently outperform the frozen baseline across multimodal tasks. Our analysis suggests that this instability arises from visual preference conflicts, where the context-agnostic nature of vision encoders induces divergent parameter updates under diverse multimodal context. To address this issue, we propose the Context-aware Visual Fine-tuning (CoVFT) framework, which explicitly incorporates multimodal context into visual adaptation. By integrating a Context Vector Extraction (CVE) and a Contextual Mixture-of-Experts (CoMoE) module, CoVFT decomposes conflicting optimization signals and enables stable, context-sensitive visual updates. Extensive experiments on 12 multimodal benchmarks demonstrate that CoVFT achieves state-of-the-art performance with superior stability. Notably, fine-tuning a 7B MLLM with CoVFT surpasses the average performance of its 13B counterpart, revealing substantial untapped potential in visual encoder optimization within MLLMs.

📄 PDF Abstract BibTeX arXiv:2603.21077

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the Loss of Context-awareness in General Instruction Fine-tuning

2024-11-05 · Yihan Wang, Andrew Bai, Nanyun Peng, Cho-Jui Hsieh

Pre-trained Large Language Models (LLMs) require post-training methods such as supervised fine-tuning (SFT) on instruction-response pairs to enable instruction following. However, this process can potentially harm existi…

BenchmarkingInstruction Following

Context-Aware Meta-Learning

2023-10-17 · Christopher Fifty, Dennis Duan, Ronald G. Junkins, Ehsan Amid 외

Large Language Models like ChatGPT demonstrate a remarkable capacity to learn new concepts during inference without any fine-tuning. However, visual models trained to detect new objects during inference have been unable …

Few-Shot Image ClassificationIn-Context LearningMeta-Learninguniversal meta-learning

Region-Level Context-Aware Multimodal Understanding

2025-08-17 · Hongliang Wei, Xianqi Zhang, Xingtao Wang, Xiaopeng Fan 외 arxiv

Despite significant progress, existing research on Multimodal Large Language Models (MLLMs) mainly focuses on general visual understanding, overlooking the ability to integrate textual context associated with objects for…

Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs

2024-04-11 · CVPR 2024 1 · Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S. Ryoo 외

Integration of Large Language Models (LLMs) into visual domain tasks, resulting in visual-LLMs (V-LLMs), has enabled exceptional performance in vision-language tasks, particularly for visual question answering (VQA). How…

DescriptiveHallucinationQuestion AnsweringSpatial Reasoning+4

OmniMem: Perturbation-aware Memory Compression for Streaming Audio-Visual LLMs

2026-05-26 · Guangzhi Sun, Yixuan Li, Yudong Yang, Chao Zhang arxiv

Audio-visual large language models (LLMs) hold strong promise for long-form video understanding, yet their long-video inference is fundamentally limited by the linear growth of video tokens and key-value (KV) caches. We …