paper-with-me

홈 › Papers

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs

2025-07-01 · Selim Kuzucu, Muhammad Ferjad Naeem, Anna Kukleva, Federico Tombari, Bernt Schiele arxiv

The integration of Large Language Model (LLMs) blocks with Vision Transformers (ViTs) holds immense promise for vision-only tasks by leveraging the rich semantic knowledge and reasoning capabilities of LLMs. However, a fundamental challenge lies in the inherent modality mismatch between text-centric pretraining of LLMs and vision-centric training of ViTs. Direct fusion often fails to fully exploit the LLM's potential and suffers from unstable finetuning. As a result, LLM blocks are kept frozen while only the vision components are learned. As a remedy to these challenges, we introduce Language-Unlocked Vision Transformers (LUViT), a novel approach that bridges this modality mismatch through a synergistic pre-training strategy. LUViT co-adapts a ViT backbone and an LLM fusion block by (1) employing Masked Auto-Encoding (MAE) to pre-train the ViT for richer visual representations, and (2) concurrently training Low-Rank Adaptation (LoRA) layers within the LLM block using the MAE objective. This joint optimization guides the ViT to produce LLM-aligned features and the LLM to effectively interpret visual information. We demonstrate through extensive experiments that LUViT significantly improves performance on various downstream vision tasks, showcasing a more effective and efficient pathway to harness LLM knowledge for visual understanding.

📄 PDF Abstract BibTeX arXiv:2507.00754

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OCC-MLLM-Alpha:Empowering Multi-modal Large Language Model for the Understanding of Occluded Objects with Self-Supervised Test-Time Learning

2024-10-02 · Shuxin Yang, Xinhan Di

There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multi-modal models fail to provide satisfactory results in describing occluded …

3D GenerationLanguage ModelingLanguage ModellingLarge Language Model+1

ForgerySleuth: Empowering Multimodal Large Language Models for Image Manipulation Detection

2024-11-29 · Zhihao Sun, Haoran Jiang, Haoran Chen, Yixin Cao 외

Multimodal large language models have unlocked new possibilities for various multimodal tasks. However, their potential in image manipulation detection remains unexplored. When directly applied to the IMD task, M-LLMs of…

Image ManipulationImage Manipulation Detection

Exploring Intrinsic Properties of Medical Images for Self-Supervised Binary Semantic Segmentation

2024-02-04 · Pranav Singh, Jacopo Cirrone

Recent advancements in self-supervised learning have unlocked the potential to harness unlabeled data for auxiliary tasks, facilitating the learning of beneficial priors. This has been particularly advantageous in fields…

DecoderImage SegmentationMedical Image AnalysisMedical Image Segmentation+3

Can GPT tell us why these images are synthesized? Empowering Multimodal Large Language Models for Forensics

2025-04-16 · Yiran He, Yun Cao, Bowen Yang, Zeyu Zhang

The rapid development of generative AI facilitates content creation and makes image manipulation easier and more difficult to detect. While multimodal Large Language Models (LLMs) have encoded rich world knowledge, they …

Few-Shot LearningImage ManipulationPrompt EngineeringWorld Knowledge

QuantMCP: Grounding Large Language Models in Verifiable Financial Reality

2025-06-07 · Yifan Zeng

Large Language Models (LLMs) hold immense promise for revolutionizing financial analysis and decision-making, yet their direct application is often hampered by issues of data hallucination and lack of access to real-time…

Decision MakingFinancial AnalysisHallucination