paper-with-me

홈 › Papers

Fusion to Enhance: Fusion Visual Encoder to Enhance Multimodal Language Model

2025-08-31 · Yifei She, Huangxuan Wu arxiv

Multimodal Large Language Models (MLLMs) have made significant progress in bridging visual perception with high-level textual reasoning. However, they face a fundamental contradiction: while excelling at complex semantic understanding, these models often fail at basic visual tasks that require precise detail perception. This deficiency primarily stems from the prevalent architectural reliance on a single vision encoder optimized for high-level semantic alignment, which inherently sacrifices the ability to capture fine-grained visual information. To address this issue, we introduce Fusion to Enhance (FtZ), a novel vision tower framework. FtZ moves beyond the single-encoder design by innovatively composing a semantically powerful anchor encoder with a perception-rich augmenting encoder via a lightweight Multi-Head Cross-Attention mechanism. Experimental results demonstrate that on several challenging benchmarks demanding fine-grained visual understanding, such as TextVQA, POPE, MMMU, MME and MM-Vet, our FtZ model significantly outperforms baselines that use only a single encoder or existing feature fusion methods. This work proves that composing heterogeneous expert encoders is an efficient and effective path to overcoming the visual perception bottleneck in current MLLMs, offering a new design paradigm for building next-generation AI systems with stronger perceptual capabilities.

📄 PDF Abstract BibTeX arXiv:2509.00664

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CMFN: Cross-Modal Fusion Network for Irregular Scene Text Recognition

2024-01-18 · Jinzhi Zheng, Ruyi Ji, Libo Zhang, Yanjun Wu 외

Scene text recognition, as a cross-modal task involving vision and text, is an important research topic in computer vision. Most existing methods use language models to extract semantic information for optimizing visual …

PositionScene Text Recognition

Effective Image Tampering Localization via Enhanced Transformer and Co-attention Fusion

2023-09-17 · Kun Guo, Haochen Zhu, Gang Cao

Powerful manipulation techniques have made digital image forgeries be easily created and widespread without leaving visual anomalies. The blind localization of tampered regions becomes quite significant for image forensi…

Image Forensics

StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation

2024-08-02

Multimodal semantic segmentation shows significant potential for enhancing segmentation accuracy in complex scenes. However, current methods often incorporate specialized feature fusion modules tailored to specific modal…

SegmentationSemantic SegmentationThermal Image Segmentation

DiffCVE: Diffusion-based Compressed Video Enhancement

2026-07-08 · Wenqiang Xiao, Wenzhuo Ma, Junxi Zhang, Zhenzhong Chen arxiv

Perceptual quality enhancement of severely compressed videos remains challenging due to complex artifact patterns and substantial information loss. Recent diffusion models have demonstrated strong generative capability f…

Video Enhancement

Audio-Visual Speech Enhancement with Score-Based Generative Models

2023-06-02 · Julius Richter, Simone Frintrop, Timo Gerkmann

This paper introduces an audio-visual speech enhancement system that leverages score-based generative models, also known as diffusion models, conditioned on visual information. In particular, we exploit audio-visual embe…

Automatic Speech RecognitionLipreadingSpeech Enhancementspeech-recognition+1