paper-with-me

홈 › Papers

IntRec: Intent-based Retrieval with Contrastive Refinement

2026-02-19 · Pourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger, Yue Lu arxiv

Retrieving user-specified objects from complex scenes remains a challenging task, especially when queries are ambiguous or involve multiple similar objects. Existing open-vocabulary detectors operate in a one-shot manner, lacking the ability to refine predictions based on user feedback. To address this, we propose IntRec, an interactive object retrieval framework that refines predictions based on user feedback. At its core is an Intent State (IS) that maintains dual memory sets for positive anchors (confirmed cues) and negative constraints (rejected hypotheses). A contrastive alignment function ranks candidate objects by maximizing similarity to positive cues while penalizing rejected ones, enabling fine-grained disambiguation in cluttered scenes. Our interactive framework provides substantial improvements in retrieval accuracy without additional supervision. On LVIS, IntRec achieves 35.4 AP, outperforming OVMR, CoDet, and CAKE by +2.3, +3.7, and +0.5, respectively. On the challenging LVIS-Ambiguous benchmark, it improves performance by +7.9 AP over its one-shot baseline after a single corrective feedback, with less than 30 ms of added latency per interaction.

📄 PDF Abstract BibTeX arXiv:2602.17639

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MVCL-DAF++: Enhancing Multimodal Intent Recognition via Prototype-Aware Contrastive Alignment and Coarse-to-Fine Dynamic Attention Fusion

2025-09-22 · Haofeng Huang, Yifei Han, Long Zhang, Bin Li 외 arxiv

Multimodal intent recognition (MMIR) suffers from weak semantic grounding and poor robustness under noisy or rare-class conditions. We propose MVCL-DAF++, which extends MVCL-DAF with two key modules: (1) Prototype-aware …

Multimodal Intent Recognition

MIntRec: A New Dataset for Multimodal Intent Recognition

2022-09-09 · Hanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou 외

Multimodal intent recognition is a significant task for understanding human language in real-world multimodal scenes. Most existing intent recognition methods have limitations in leveraging the multimodal information due…

Intent RecognitionMultimodal Intent Recognition

Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection

2025-08-22 · Ankan Mullick, Saransh Sharma, Abhik Jana, Pawan Goyal arxiv

The rise of multimodal data, integrating text, audio, and visuals, has created new opportunities for studying multimodal tasks such as intent detection. This work investigates the effectiveness of Large Language Models (…

Intent Detection

MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations

2024-03-16 · Hanlei Zhang, Xin Wang, Hua Xu, Qianrui Zhou 외

Multimodal intent recognition poses significant challenges, requiring the incorporation of non-verbal modalities from real-world contexts to enhance the comprehension of human intentions. Existing benchmark datasets are …

Intent RecognitionMultimodal Intent Recognition

IntPro: A Proxy Agent for Context-Aware Intent Understanding via Retrieval-conditioned Inference

2026-02-10 · Guanming Liu, Meng Wu, Peng Zhang, Yu Zhang 외 arxiv

Large language models (LLMs) have become integral to modern Human-AI collaboration workflows, where accurately understanding user intent serves as a crucial step for generating satisfactory responses. Context-aware inten…