paper-with-me

홈 › Papers

Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions

2024-02-17 · Wenxuan Wang, Yisi Zhang, Xingjian He, Yichen Yan, Zijia Zhao, Xinlong Wang, Jing Liu

Visual grounding (VG) aims at locating the foreground entities that match the given natural language expressions. Previous datasets and methods for classic VG task mainly rely on the prior assumption that the given expression must literally refer to the target object, which greatly impedes the practical deployment of agents in real-world scenarios. Since users usually prefer to provide intention-based expression for the desired object instead of covering all the details, it is necessary for the agents to interpret the intention-driven instructions. Thus, in this work, we take a step further to the intention-driven visual-language (V-L) understanding. To promote classic VG towards human intention interpretation, we propose a new intention-driven visual grounding (IVG) task and build a large-scale IVG dataset termed IntentionVG with free-form intention expressions. Considering that practical agents need to move and find specific targets among various scenarios to realize the grounding task, our IVG task and IntentionVG dataset have taken the crucial properties of both multi-scenario perception and egocentric view into consideration. Besides, various types of models are set up as the baselines to realize our IVG task. Extensive experiments on our IntentionVG dataset and baselines demonstrate the necessity and efficacy of our method for the V-L field. To foster future research in this direction, our newly built dataset and baselines will be publicly available at https://github.com/Rubics-Xuan/IVG.

📄 PDF Abstract BibTeX arXiv:2402.11265

Code (1)

rubics-xuan/ivg 공식 구현

Tasks

Visual Grounding

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

HeGeL: A Novel Dataset for Geo-Location from Hebrew Text

2023-07-02 · Tzuf Paz-Argaman, Tal Bauman, Itai Mondshine, Itzhak Omer 외

The task of textual geolocation - retrieving the coordinates of a place based on a free-form language description - calls for not only grounding but also natural language understanding and geospatial reasoning. Even thou…

Natural Language UnderstandingRetrieval

Beyond the Literal: Decomposing Pragmatic Intent in Multimodal Meme Understanding

2026-06-02 · Zhengyi Zhao, Shubo Zhang, Zezhong Wang, Luyao Ye 외 arxiv

When asked what a meme or sarcastic post means, Large Vision Language Models (LVLMs) tend to describe what the image shows rather than what the author is trying to communicate. Standard instruction tuning entangles a pos…

Towards Task Understanding in Visual Settings

2018-11-28 · Sebastin Santy, Wazeer Zulfikar, Rishabh Mehrotra, Emine Yilmaz

We consider the problem of understanding real world tasks depicted in visual images. While most existing image captioning methods excel in producing natural language descriptions of visual scenes involving human tasks, t…

Image CaptioningText Generation

ViPE: Visualise Pretty-much Everything

2023-10-16 · Hassan Shahmohammadi, Adhiraj Ghosh, Hendrik P. A. Lensch

Figurative and non-literal expressions are profoundly integrated in human communication. Visualising such expressions allow us to convey our creative thoughts, and evoke nuanced emotions. Recent text-to-image models like…

Caption GenerationFigurative Language VisualizationImage Generation

Beyond Bounding Box: Multimodal Knowledge Learning for Object Detection

2022-05-09 · Weixin Feng, Xingyuan Bu, Chenchen Zhang, Xubin Li

Multimodal supervision has achieved promising results in many visual language understanding tasks, where the language plays an essential role as a hint or context for recognizing and locating instances. However, due to t…

Objectobject-detectionObject Detection