paper-with-me

홈 › Papers

Referring Multiple Regions with Large Multimodal Models via Contextual Latent Steering

2026-05-03 · Yun Xing, Hanyuan Liu, Jiahao Nie, Shijian Lu arxiv

Large Multimodal Models (LMMs) have recently demonstrated their proficiency in holistic visual comprehension. However, most of them struggle to tackle region-level perception guided by visual prompts, especially for cases where multiple regions are referred simultaneously, or scenarios where global contexts are necessary for precise visual referring. We introduce Contextual Latent Steering (CSteer), a training-free approach for guiding general LMMs to refer multiple regions contextually, without expensive fine-tuning or architectural modifications. CSteer starts with pre-computing contextual vectors that implicitly represent visual referring behaviors, such as differentiation among regions and attention to global contexts, followed by representation editing during inference time. Experimental results on multiple datasets indicate that general LMMs with CSteer outperform tailored referring LMMs in most cases, suggesting a promising solution in training-free, and setting new state-of-the-art for this field. Code is available at https://github.com/xing0047/csteer.git.

📄 PDF Abstract BibTeX arXiv:2605.01827

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

XeMap: Contextual Referring in Large-Scale Remote Sensing Environments

2025-04-30 · Yuxi Li, Lu Si, Yujie Hou, Chengaung Liu 외

Advancements in remote sensing (RS) imagery have provided high-resolution detail and vast coverage, yet existing methods, such as image-level captioning/retrieval and object-level detection/segmentation, often fail to ca…

ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models

2024-07-31 · Mingrui Wu, Xinyue Cai, Jiayi Ji, Jiale Li 외

In this work, we propose a training-free method to inject visual prompts into Multimodal Large Language Models (MLLMs) through test-time optimization of a learnable latent variable. We observe that attention, as the core…

Domain GeneralizationPrompt Learning

EAGLE: Towards Efficient Arbitrary Referring Visual Prompts Comprehension for Multimodal Large Language Models

2024-09-25 · Jiacheng Zhang, Yang Jiao, Shaoxiang Chen, Jingjing Chen 외

Recently, Multimodal Large Language Models (MLLMs) have sparked great research interests owing to their exceptional content-reasoning and instruction-following capabilities. To effectively instruct an MLLM, in addition t…

Instruction Following

DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution

2025-01-01 · CVPR 2025 1 · Yuzhong Zhao, Feng Liu, Yue Liu, Mingxiang Liao 외

One important task of multimodal models is to translate referred image regions to human preferred language descriptions. Existing methods, however, ignore the resolution adaptability needs of different tasks, which h…

Attribute

Referring Transformer: A One-step Approach to Multi-task Visual Grounding

2021-06-06 · NeurIPS 2021 12 · Muchen Li, Leonid Sigal

As an important step towards visual reasoning, visual grounding (e.g., phrase localization, referring expression comprehension/segmentation) has been widely explored Previous approaches to referring expression comprehens…

DecoderReferring ExpressionReferring Expression ComprehensionReferring Expression Segmentation+3