paper-with-me

홈 › Papers

V?: Guided Visual Search as a Core Mechanism in Multimodal LLMs

2024-01-01 · CVPR 2024 1 · Penghao Wu, Saining Xie

When we look around and perform complex tasks how we see and selectively process what we see is crucial. However the lack of this visual search mechanism in current multimodal LLMs (MLLMs) hinders their ability to focus on important visual details especially when handling high-resolution and visually crowded images. To address this we introduce V* an LLM-guided visual search mechanism that employs the world knowledge in LLMs for efficient visual querying. When combined with an MLLM this mechanism enhances collaborative reasoning contextual understanding and precise visual grounding. This integration results in a new MLLM meta-architecture named Show sEArch and TelL (SEAL). We further create V*Bench a benchmark specifically designed to evaluate MLLMs in their ability to process high-resolution images and focus on visual details. Our study highlights the necessity of incorporating visual search capabilities into multimodal systems. The code is available at https://github.com/penghao-wu/vstar

📄 PDF Abstract BibTeX

Code (1)

penghao-wu/vstar 공식 구현 pytorch

Tasks

Visual GroundingWorld Knowledge

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

2023-12-21 · Penghao Wu, Saining Xie

When we look around and perform complex tasks, how we see and selectively process what we see is crucial. However, the lack of this visual search mechanism in current multimodal LLMs (MLLMs) hinders their ability to focu…

Visual Question AnsweringWorld Knowledge

LapsCore: Language-Guided Person Search via Color Reasoning

2021-01-01 · ICCV 2021 10 · Yushuang Wu, Zizheng Yan, Xiaoguang Han, Guanbin Li 외

The key point of language-guided person search is to construct the cross-modal association between visual and textual input. Existing methods focus on designing multimodal attention mechanisms and novel cross-modal l…

ColorizationImage ColorizationPerson SearchRepresentation Learning

From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning

2026-03-04 · Ruilin Luo, Chufan Shi, Yizhen Zhang, Cheng Yang 외 arxiv

The cold-start initialization stage plays a pivotal role in training Multimodal Large Reasoning Models (MLRMs), yet its mechanisms remain insufficiently understood. To analyze this stage, we introduce the Visual Attentio…

Multimodal Reasoning

MI-Pruner: Crossmodal Mutual Information-guided Token Pruner for Efficient MLLMs

2026-04-03 · Jiameng Li, Aleksei Tiulpin, Matthew B. Blaschko arxiv

For multimodal large language models (MLLMs), visual information is relatively sparse compared with text. As a result, research on visual pruning emerges for efficient inference. Current approaches typically measure toke…

Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation

2026-05-05 · Quanxing Xu, Ling Zhou, Xian Zhong, Xiaohua Huang 외 arxiv

With advances in multimodal research and deep learning, Multimodal Large Language Models (MLLMs) have emerged as a powerful paradigm for a wide range of multimodal tasks. As a core problem in vision-language research, Vi…

Visual Question Answering