paper-with-me

Papers

Weakly Supervised Attention Learning for Textual Phrases Grounding

2018-05-01 · Zhiyuan Fang, Shu Kong, Tianshu Yu, Yezhou Yang

Grounding textual phrases in visual content is a meaningful yet challenging problem with various potential applications such as image-text inference or text-driven multimedia interaction. Most of the current existing methods adopt the supervised learning mechanism which requires ground-truth at pixel level during training. However, fine-grained level ground-truth annotation is quite time-consuming and severely narrows the scope for more general applications. In this extended abstract, we explore methods to localize flexibly image regions from the top-down signal (in a form of one-hot label or natural languages) with a weakly supervised attention learning mechanism. In our model, two types of modules are utilized: a backbone module for visual feature capturing, and an attentive module generating maps based on regularized bilinear pooling. We construct the model in an end-to-end fashion which is trained by encouraging the spatial attentive map to shift and focus on the region that consists of the best matched visual features with the top-down signal. We demonstrate the preliminary yet promising results on a testbed that is synthesized with multi-label MNIST data.

📄 PDF Abstract BibTeX arXiv:1805.00545

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Weakly-supervised Visual Grounding of Phrases with Linguistic Structures

2017-05-03 · CVPR 2017 7 · Fanyi Xiao, Leonid Sigal, Yong Jae Lee

We propose a weakly-supervised approach that takes image-sentence pairs as input and learns to visually ground (i.e., localize) arbitrary linguistic phrases, in the form of spatial attention masks. Specifically, the mode…

SentenceVisual Grounding

Weak Supervision and Referring Attention for Temporal-Textual Association Learning

2020-06-21 · Zhiyuan Fang, Shu Kong, Zhe Wang, Charless Fowlkes 외

A system capturing the association between video frames and textual queries offer great potential for better video analysis. However, training such a system in a fully supervised way inevitably demands a meticulously cur…

Weakly-Supervised Visual-Textual Grounding with Semantic Prior Refinement

2023-05-18 · Davide Rigoni, Luca Parolari, Luciano Serafini, Alessandro Sperduti 외

Using only image-sentence pairs, weakly-supervised visual-textual grounding aims to learn region-phrase correspondences of the respective entity mentions. Compared to the supervised approach, learning is more difficult s…

Sentence

Finding "It": Weakly-Supervised Reference-Aware Visual Grounding in Instructional Videos

2018-06-01 · CVPR 2018 6 · De-An Huang, Shyamal Buch, Lucio Dery, Animesh Garg 외

Grounding textual phrases in visual content with standalone image-sentence pairs is a challenging task. When we consider grounding in instructional videos, this problem becomes profoundly more complex: the latent tempora…

Multiple Instance LearningSentenceVisual Grounding

Grounding of Textual Phrases in Images by Reconstruction

2015-11-12 · Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell 외

Grounding (i.e. localizing) arbitrary, free-form textual phrases in visual content is a challenging problem with many applications for human-computer interaction and image-text reference resolution. Few datasets provide …

Language ModelingLanguage ModellingNatural Language Visual GroundingPhrase Grounding+1