paper-with-me

Papers

Compositional Temporal Visual Grounding of Natural Language Event Descriptions

2019-12-04 · Jonathan C. Stroud, Ryan McCaffrey, Rada Mihalcea, Jia Deng, Olga Russakovsky

Temporal grounding entails establishing a correspondence between natural language event descriptions and their visual depictions. Compositional modeling becomes central: we first ground atomic descriptions "girl eating an apple," "batter hitting the ball" to short video segments, and then establish the temporal relationships between the segments. This compositional structure enables models to recognize a wider variety of events not seen during training through recognizing their atomic sub-events. Explicit temporal modeling accounts for a wide variety of temporal relationships that can be expressed in language: e.g., in the description "girl stands up from the table after eating an apple" the visual ordering of the events is reversed, with first "eating an apple" followed by "standing up from the table." We leverage these observations to develop a unified deep architecture, CTG-Net, to perform temporal grounding of natural language event descriptions to videos. We demonstrate that our system outperforms prior state-of-the-art methods on the DiDeMo, Tempo-TL, and Tempo-HL temporal grounding datasets.

📄 PDF Abstract BibTeX arXiv:1912.02256

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning

2022-03-24 · CVPR 2022 1 · Juncheng Li, Junlin Xie, Long Qian, Linchao Zhu 외

Temporal grounding in videos aims to localize one target video segment that semantically corresponds to a given query sentence. Thanks to the semantic diversity of natural language descriptions, temporal grounding allows…

DiversitySemantic correspondenceSentence

Differentiable Parsing and Visual Grounding of Natural Language Instructions for Object Placement

2022-10-01 · Zirui Zhao, Wee Sun Lee, David Hsu

We present a new method, PARsing And visual GrOuNding (ParaGon), for grounding natural language in object placement tasks. Natural language generally describes objects and spatial relations with compositionality and ambi…

Graph Neural NetworkObjectRelational ReasoningVisual Grounding

Variational Cross-Graph Reasoning and Adaptive Structured Semantics Learning for Compositional Temporal Grounding

2023-01-22 · Juncheng Li, Siliang Tang, Linchao Zhu, Wenqiao Zhang 외

Temporal grounding is the task of locating a specific segment from an untrimmed video according to a query sentence. This task has achieved significant momentum in the computer vision community as it enables activity gro…

DiversitySemantic correspondenceSentence

Learning to Compose and Reason with Language Tree Structures for Visual Grounding

2019-06-05 · Richang Hong, Daqing Liu, Xiaoyu Mo, Xiangnan He 외

Grounding natural language in images, such as localizing "the black dog on the left of the tree", is one of the core problems in artificial intelligence, as it needs to comprehend the fine-grained and compositional langu…

Visual GroundingVisual Reasoning

Exploiting Temporal Relationships in Video Moment Localization with Natural Language

2019-08-11 · Songyang Zhang, Jinsong Su, Jiebo Luo

We address the problem of video moment localization with natural language, i.e. localizing a video segment described by a natural language sentence. While most prior work focuses on grounding the query as a whole, tempor…

Sentence