paper-with-me

홈 › Papers

Improving Pre-trained Vision-and-Language Embeddings for Phrase Grounding

2021-11-01 · EMNLP 2021 11 · Zi-Yi Dou, Nanyun Peng

Phrase grounding aims to map textual phrases to their associated image regions, which can be a prerequisite for multimodal reasoning and can benefit tasks requiring identifying objects based on language. With pre-trained vision-and-language models achieving impressive performance across tasks, it remains unclear if we can directly utilize their learned embeddings for phrase grounding without fine-tuning. To this end, we propose a method to extract matched phrase-region pairs from pre-trained vision-and-language embeddings and propose four fine-tuning objectives to improve the model phrase grounding ability using image-caption data without any supervised grounding signals. Experiments on two representative datasets demonstrate the effectiveness of our objectives, outperforming baseline models in both weakly-supervised and supervised phrase grounding settings. In addition, we evaluate the aligned embeddings on several other downstream tasks and show that we can achieve better phrase grounding without sacrificing representation generality.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningPhrase Grounding

Similar Papers 제목 키워드 기반

Contrastive Learning for Weakly Supervised Phrase Grounding

2020-06-17 · ECCV 2020 8 · Tanmay Gupta, Arash Vahdat, Gal Chechik, Xiaodong Yang 외

Phrase grounding, the problem of associating image regions to caption words, is a crucial component of vision-language tasks. We show that phrase grounding can be learned by optimizing word-region attention to maximize a…

Contrastive LearningLanguage ModelingLanguage ModellingPhrase Grounding

Grounding of Textual Phrases in Images by Reconstruction

2015-11-12 · Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell 외

Grounding (i.e. localizing) arbitrary, free-form textual phrases in visual content is a challenging problem with many applications for human-computer interaction and image-text reference resolution. Few datasets provide …

Language ModelingLanguage ModellingNatural Language Visual GroundingPhrase Grounding+1

Catalog Phrase Grounding (CPG): Grounding of Product Textual Attributes in Product Images for e-commerce Vision-Language Applications

2023-08-30 · Wenyi Wu, Karim Bouyarmane, Ismail Tutar

We present Catalog Phrase Grounding (CPG), a model that can associate product textual data (title, brands) into corresponding regions of product images (isolated product region, brand logo region) for e-commerce vision-l…

Decoderobject-detectionObject DetectionPhrase Grounding

Conditional Image-Text Embedding Networks

2017-11-22 · ECCV 2018 9 · Bryan A. Plummer, Paige Kordas, M. Hadi Kiapour, Shuai Zheng 외

This paper presents an approach for grounding phrases in images which jointly learns multiple text-conditioned embeddings in a single end-to-end model. In order to differentiate text phrases into semantically distinct su…

Phrase Grounding

Utilizing Every Image Object for Semi-supervised Phrase Grounding

2020-11-05 · Haidong Zhu, Arka Sadhu, Zhaoheng Zheng, Ram Nevatia

Phrase grounding models localize an object in the image given a referring expression. The annotated language queries available during training are limited, which also limits the variations of language combinations that a…

Phrase GroundingReferring Expression