paper-with-me

홈 › Papers

Disentangled Motif-aware Graph Learning for Phrase Grounding

2021-04-13 · Zongshen Mu, Siliang Tang, Jie Tan, Qiang Yu, Yueting Zhuang

In this paper, we propose a novel graph learning framework for phrase grounding in the image. Developing from the sequential to the dense graph model, existing works capture coarse-grained context but fail to distinguish the diversity of context among phrases and image regions. In contrast, we pay special attention to different motifs implied in the context of the scene graph and devise the disentangled graph network to integrate the motif-aware contextual information into representations. Besides, we adopt interventional strategies at the feature and the structure levels to consolidate and generalize representations. Finally, the cross-modal attention network is utilized to fuse intra-modal features, where each phrase can be computed similarity with regions to select the best-grounded one. We validate the efficiency of disentangled and interventional graph network (DIGN) through a series of ablation studies, and our model achieves state-of-the-art performance on Flickr30K Entities and ReferIt Game benchmarks.

📄 PDF Abstract BibTeX arXiv:2104.06008

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityGraph LearningPhrase Grounding

Similar Papers 제목 키워드 기반

Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal Grounding

2025-10-23 · Minseok Kang, Minhyeok Lee, Minjung Kim, Donghyeong Kim 외 arxiv

Video Temporal Grounding (VTG) aims to localize temporal segments in long, untrimmed videos that align with a given natural language query. This task typically comprises two subtasks: Moment Retrieval (MR) and Highlight …

Highlight DetectionMoment RetrievalVideo Grounding

Learning Cross-modal Context Graph for Visual Grounding

2020-02-13 · AAAI-2020 2020 2 · Yongfei Liu; Bo Wan; Xiaodan Zhu; Xuming He

Visual grounding is a ubiquitous building block in many vision-language tasks and yet remains challenging due to large variations in visual and linguistic features of grounding entities, strong context effect and the res…

Graph MatchingGraph Neural NetworkLanguage ModellingNatural Language Visual Grounding+2

Learning Cross-modal Context Graph for Visual Grounding

2019-11-20 · Yongfei Liu, Bo Wan, Xiaodan Zhu, Xuming He

Visual grounding is a ubiquitous building block in many vision-language tasks and yet remains challenging due to large variations in visual and linguistic features of grounding entities, strong context effect and the res…

Graph MatchingGraph Neural NetworkVisual Grounding

Position-aware Location Regression Network for Temporal Video Grounding

2022-04-12 · Sunoh Kim, Kimin Yun, Jin Young Choi

The key to successful grounding for video surveillance is to understand a semantic phrase corresponding to important actors and objects. Conventional methods ignore comprehensive contexts for the phrase or require heavy …

PositionregressionVideo Grounding

Motif-aware Riemannian Graph Neural Network with Generative-Contrastive Learning

2024-01-02 · Li Sun, Zhenhao Huang, Zixi Wang, Feiyang Wang 외

Graphs are typical non-Euclidean data of complex structures. In recent years, Riemannian graph representation learning has emerged as an exciting alternative to Euclidean ones. However, Riemannian methods are still in an…

Contrastive LearningGraph Neural NetworkGraph Representation LearningRepresentation Learning