paper-with-me

홈 › Papers

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models

2025-08-09 · Jianting Tang, Yubo Wang, Haoyu Cao, Linli Xu arxiv

Mainstream Multimodal Large Language Models (MLLMs) achieve visual understanding by using a vision projector to bridge well-pretrained vision encoders and large language models (LLMs). The inherent gap between visual and textual modalities makes the embeddings from the vision projector critical for visual comprehension. However, current alignment approaches treat visual embeddings as contextual cues and merely apply auto-regressive supervision to textual outputs, neglecting the necessity of introducing equivalent direct visual supervision, which hinders the potential finer alignment of visual embeddings. In this paper, based on our analysis of the refinement process of visual embeddings in the LLM's shallow layers, we propose BASIC, a method that utilizes refined visual embeddings within the LLM as supervision to directly guide the projector in generating initial visual embeddings. Specifically, the guidance is conducted from two perspectives: (i) optimizing embedding directions by reducing angles between initial and supervisory embeddings in semantic space; (ii) improving semantic matching by minimizing disparities between the logit distributions of both visual embeddings. Without additional supervisory models or artificial annotations, BASIC significantly improves the performance of MLLMs across a wide range of benchmarks, demonstrating the effectiveness of our introduced direct visual supervision.

📄 PDF Abstract BibTeX arXiv:2508.06895

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning

2025-08-06 · Jianghangfan Zhang, Yibo Yan, Kening Zheng, Xin Zou 외 arxiv

Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities but often struggle with complex, multi-step mathematical reasoning, where minor errors in visual perception or logical deduction can lead to co…

Mathematical Reasoning

StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues

2026-02-23 · Zanxi Ruan, Songqun Gao, Qiuyu Kong, Yiming Wang 외 arxiv

Edge-based representations are fundamental cues for visual understanding, a principle rooted in early vision research and still central today. We extend this principle to vision-language alignment, showing that isolating…

Cross-Modal Retrieval

Boosting Transition-based AMR Parsing with Refined Actions and Auxiliary Analyzers

2015-07-01 · IJCNLP 2015 7 · Chuan Wang, Nianwen Xue, Sameer Pradhan
AMR ParsingDependency Parsing

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

2025-03-27 · Dian Zheng, Ziqi Huang, Hongbo Liu, Kai Zou 외

Video generation has advanced significantly, evolving from producing unrealistic outputs to generating videos that appear visually convincing and temporally coherent. To evaluate these video generative models, benchmarks…

Anomaly DetectionVideo Generation

Boosting Generalizable Depth Estimation in Endoscopy by Mixture of Lightweight Experts and Intrinsic Image Alignment

2026-08-01 · Liangjing Shao, Beilei Cui, Yiming Huang, Changjing Liu 외 arxiv

Depth estimation is a significant task for 3D perception in endoscopic surgeries. However, illumination interference and feature diversity in various endoscopic scenes are still challenges for generalizable depth estimat…

parameter-efficient fine-tuningDepth Estimation