paper-with-me

Papers

CoNAN: A Complementary Neighboring-based Attention Network for Referring Expression Generation

2020-12-01 · COLING 2020 8 · Jungjun Kim, Hanbin Ko, Jialin Wu

Daily scenes are complex in the real world due to occlusion, undesired lighting conditions, etc. Although humans handle those complicated environments well, they evoke challenges for machine learning systems to identify and describe the target without ambiguity. Most previous research focuses on mining discriminating features within the same category for the target object. One the other hand, as the scene becomes more complicated, human frequently uses the neighbor objects as complementary information to describe the target one. Motivated by that, we propose a novel Complementary Neighboring-based Attention Network (CoNAN) that explicitly utilizes the visual differences between the target object and its highly-related neighbors. These highly-related neighbors are determined by an attentional ranking module, as complementary features, highlighting the discriminating aspects for the target object. The speaker module then takes the visual difference features as an additional input to generate the expression. Our qualitative and quantitative results on the dataset RefCOCO, RefCOCO+, and RefCOCOg demonstrate that our generated expressions outperform other state-of-the-art models by a clear margin.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectReferring ExpressionReferring expression generation

Similar Papers 제목 키워드 기반

Improving Referring Expression Grounding with Cross-modal Attention-guided Erasing

2019-03-03 · CVPR 2019 6 · Xihui Liu, ZiHao Wang, Jing Shao, Xiaogang Wang 외

Referring expression grounding aims at locating certain objects or persons in an image with a referring expression, where the key challenge is to comprehend and align various types of information from visual and textual …

Referring Expression

CONAN: Complementary Pattern Augmentation for Rare Disease Detection

2019-11-26 · Limeng Cui, Siddharth Biswal, Lucas M. Glass, Greg Lever 외

Rare diseases affect hundreds of millions of people worldwide but are hard to detect since they have extremely low prevalence rates (varying from 1/1,000 to 1/200,000 patients) and are massively underdiagnosed. How do we…

Neighbourhood Watch: Referring Expression Comprehension via Language-guided Graph Attention Networks

2018-12-12 · CVPR 2019 6 · Peng Wang, Qi Wu, Jiewei Cao, Chunhua Shen 외

The task in referring expression comprehension is to localise the object instance in an image described by a referring expression phrased in natural language. As a language-to-vision matching task, the key to this proble…

Graph AttentionObjectReferring ExpressionReferring Expression Comprehension

Expression Prompt Collaboration Transformer for Universal Referring Video Object Segmentation

2023-08-08 · Jiajun Chen, Jiacheng Lin, Guojin Zhong, Haolong Fu 외

Audio-guided Video Object Segmentation (A-VOS) and Referring Video Object Segmentation (R-VOS) are two highly related tasks that both aim to segment specific objects from video sequences according to expression prompts. …

Contrastive LearningObjectReferring Expression SegmentationReferring Video Object Segmentation+4

Searching for Ambiguous Objects in Videos using Relational Referring Expressions

2019-08-03 · Hazan Anayurt, Sezai Artun Ozyegin, Ulfet Cetin, Utku Aktas 외

Humans frequently use referring (identifying) expressions to refer to objects. Especially in ambiguous settings, humans prefer expressions (called relational referring expressions) that describe an object with respect to…

Deep AttentionNatural Language Visual GroundingObjectReferring Expression