paper-with-me

홈 › Papers

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval

2024-12-15 · Zelong Sun, Dong Jing, Guoxing Yang, Nanyi Fei, Zhiwu Lu

Composed Image Retrieval (CIR) aims to retrieve target images from candidate set using a hybrid-modality query consisting of a reference image and a relative caption that describes the user intent. Recent studies attempt to utilize Vision-Language Pre-training Models (VLPMs) with various fusion strategies for addressing the task.However, these methods typically fail to simultaneously meet two key requirements of CIR: comprehensively extracting visual information and faithfully following the user intent. In this work, we propose CIR-LVLM, a novel framework that leverages the large vision-language model (LVLM) as the powerful user intent-aware encoder to better meet these requirements. Our motivation is to explore the advanced reasoning and instruction-following capabilities of LVLM for accurately understanding and responding the user intent. Furthermore, we design a novel hybrid intent instruction module to provide explicit intent guidance at two levels: (1) The task prompt clarifies the task requirement and assists the model in discerning user intent at the task level. (2) The instance-specific soft prompt, which is adaptively selected from the learnable prompt pool, enables the model to better comprehend the user intent at the instance level compared to a universal prompt for all instances. CIR-LVLM achieves state-of-the-art performance across three prominent benchmarks with acceptable inference efficiency. We believe this study provides fundamental insights into CIR-related fields.

📄 PDF Abstract BibTeX arXiv:2412.11087

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalInstruction FollowingLanguage ModelingLanguage ModellingRetrieval

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

LIT: Large Language Model Driven Intention Tracking for Proactive Human-Robot Collaboration -- A Robot Sous-Chef Application

2024-06-19 · Zhe Huang, John Pohovey, Ananya Yammanuru, Katherine Driggs-Campbell

Large Language Models (LLM) and Vision Language Models (VLM) enable robots to ground natural language prompts into control actions to achieve tasks in an open world. However, when applied to a long-horizon collaborative …

Language ModelingLanguage ModellingLarge Language Model

REVISION:Reflective Intent Mining and Online Reasoning Auxiliary for E-commerce Visual Search System Optimization

2025-10-26 · Yiwen Tang, Qiuyu Zhao, Zenghui Sun, Jinsong Lan 외 arxiv

In Taobao e-commerce visual search, user behavior analysis reveals a large proportion of no-click requests, suggesting diverse and implicit user intents. These intents are expressed in various forms and are difficult to …

Generic Intent Representation in Web Search

2019-07-24 · Hongfei Zhang, Xia Song, Chenyan Xiong, Corby Rosset 외

This paper presents GEneric iNtent Encoder (GEN Encoder) which learns a distributed representation space for user intent in search. Leveraging large scale user clicks from Bing search logs as weak supervision of user int…

Multi-Task Learning

A scalable framework for learning from implicit user feedback to improve natural language understanding in large-scale conversational AI systems

2020-10-23 · EMNLP 2021 11 · Sunghyun Park, Han Li, Ameen Patel, Sidharth Mudgal 외

Natural Language Understanding (NLU) is an established component within a conversational AI or digital assistant system, and it is responsible for producing semantic understanding of a user request. We propose a scalable…

Natural Language Understanding

EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning

2026-05-17 · Zeyu Wang, Chang Liu, Eduardus Tjitrahardja, Yuntao Wang 외 arxiv

Despite extensive efforts on egocentric video datasets and benchmarks, understanding users' internal states, which is crucial for enabling seamless AI assistant experiences, remains largely overlooked. In this work, we i…