paper-with-me

홈 › Papers

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception

2026-06-17 · Chengzhi Mao, Xudong Lin, Wen-Sheng Chu arxiv

Vision foundation models are typically trained as static feature extractors, placing the burden of task adaptation onto large downstream models. We propose an alternative paradigm: instead of solely feeding visual features into language models, we use language itself to dynamically guide the vision encoder. Our method, Language-Instructed Vision Embeddings (LIVE), leverages language as high-level guidance to produce task-centric embeddings at inference time, removing the need for task-specific retraining. This enables the encoder to focus on contextually relevant aspects of the input, yielding more controllable and generalizable representations. Empirically, LIVE reduces visual hallucinations (+34 points on MMVP), surpasses vision-language models with orders of magnitude more parameters on visual question answering, and generalizes to unseen instructions and tasks -- offering a direct path toward adaptive, instruction-driven visual intelligence.

📄 PDF Abstract BibTeX arXiv:2606.19584

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

MFVLR: Multi-domain Fine-grained Vision-Language Reconstruction for Generalizable Diffusion Face Forgery Detection and Localization

2026-05-11 · Yaning Zhang, Tianyi Wang, Zan Gao, Yibo Zhao 외 arxiv

The swift advancement in photo-realistic face generation technology has sparked considerable concerns across society and academia, emphasizing the requirement of generalizable face forgery detection and localization meth…

Representation Learning

SAGE: Bridging Semantic and Actionable Parts for GEneralizable Manipulation of Articulated Objects

2023-12-03 · Haoran Geng, Songlin Wei, Congyue Deng, Bokui Shen 외

To interact with daily-life articulated objects of diverse structures and functionalities, understanding the object parts plays a central role in both user instruction comprehension and task execution. However, the possi…

Language ModellingObject

Driver Activity Classification Using Generalizable Representations from Vision-Language Models

2024-04-23 · Ross Greer, Mathias Viborg Andersen, Andreas Møgelmose, Mohan Trivedi

Driver activity classification is crucial for ensuring road safety, with applications ranging from driver assistance systems to autonomous vehicle control transitions. In this paper, we present a novel approach leveragin…

Action Recognition

Coffee: Controllable Diffusion Fine-tuning

2025-11-18 · Ziyao Zeng, Jingcheng Ni, Ruyi Liu, Alex Wong arxiv

Text-to-image diffusion models can generate diverse content with flexible prompts, which makes them well-suited for customization through fine-tuning with a small amount of user-provided data. However, controllable fine-…

Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features

2024-10-16 · Makram Chahine, Alex Quach, Alaa Maalouf, Tsun-Hsuan Wang 외

End-to-end learning directly maps sensory inputs to actions, creating highly integrated and efficient policies for complex robotics tasks. However, such models often struggle to generalize beyond their training scenarios…

Visual Navigation