paper-with-me

Papers

FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions

2026-03-18 · Peisen Zhao, Xiaopeng Zhang, Mingxing Xu, Ruoyu Sun, Zewei Du, Dunzheng Wang, Guanghao Zheng, Haohang Xu, Zhibo Zhang, Yuhang Zhang, Yi Ai, Lin Liu, Qi Tian arxiv

While Multimodal Large Language Models (MLLMs) have experienced rapid advancements, their visual encoders frequently remain a performance bottleneck. Conventional CLIP-based encoders struggle with dense spatial tasks due to the loss of visual details caused by low-resolution pretraining and the reliance on noisy, coarse web-crawled image-text pairs. To overcome these limitations, we introduce FineViT, a novel vision encoder specifically designed to unlock fine-grained perception. By replacing coarse web data with dense recaptions, we systematically mitigate information loss through a progressive training paradigm.: first, the encoder is trained from scratch at a high native resolution on billions of global recaptioned image-text pairs, establishing a robust, detail rich semantic foundation. Subsequently, we further enhance its local perception through LLM alignment, utilizing our curated FineCap-450M dataset that comprises over $450$ million high quality local captions. Extensive experiments validate the effectiveness of the progressive strategy. FineViT achieves state-of-the-art zero-shot recognition and retrieval performance, especially in long-context retrieval, and consistently outperforms multimodal visual encoders such as SigLIP2 and Qwen-ViT when integrated into MLLMs. We hope FineViT could serve as a powerful new baseline for fine-grained visual perception.

📄 PDF Abstract BibTeX arXiv:2603.17326

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model

2025-12-14 · Zhou Tao, Shida Wang, Yongxiang Hua, Haoyu Cao 외 arxiv

Multimodal Large Language Models have achieved impressive performance on a variety of vision-language tasks, yet their fine-grained visual perception and precise spatial reasoning remain limited. In this work, we introdu…

Spatial ReasoningVisual Reasoning

Fine-grained spatial-temporal perception for gas leak segmentation

2025-05-01 · Xinlong Zhao, Shan Du

Gas leaks pose significant risks to human health and the environment. Despite long-standing concerns, there are limited methods that can efficiently and accurately detect and segment leaks due to their concealed appearan…

DecoderSegmentation

SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning

2026-02-07 · Yancheng Long, Yankai Yang, Hongyang Wei, Wei Chen 외 arxiv

Online Reinforcement Learning (RL) offers a promising avenue for complex image editing but is currently constrained by the scarcity of reliable and fine-grained reward signals. Existing evaluators frequently struggle wit…

Reinforcement LearningSpatial ReasoningImage Editing

KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering

2026-01-14 · Zhiyang Li, Ao Ke, Yukun Cao, Xike Xie arxiv

Multi-modal Large Language Models (MLLMs) for Visual Question Answering (VQA) often suffer from dual limitations: knowledge hallucination and insufficient fine-grained visual perception. Crucially, we identify that commo…

Visual Question Answering

FaVChat: Unlocking Fine-Grained Facail Video Understanding with Multimodal Large Language Models

2025-03-12 · Fufangchen Zhao, Ming Li, Linrui Xu, Wenhao Jiang 외

Video-based multimodal large language models (VMLLMs) have demonstrated remarkable potential in cross-modal video understanding. However, their abilities in fine-grained face comprehension remain largely underexplored. G…

Mixture-of-ExpertsQuestion AnsweringVideo SummarizationVideo Understanding