paper-with-me

Papers

Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding

2023-01-01 · CVPR 2023 1 · Chaolei Tan, Zihang Lin, Jian-Fang Hu, Wei-Shi Zheng, JianHuang Lai

Video Paragraph Grounding (VPG) is an essential yet challenging task in vision-language understanding, which aims to jointly localize multiple events from an untrimmed video with a paragraph query description. One of the critical challenges in addressing this problem is to comprehend the complex semantic relations between visual and textual modalities. Previous methods focus on modeling the contextual information between the video and text from a single-level perspective (i.e., the sentence level), ignoring rich visual-textual correspondence relations at different semantic levels, e.g., the video-word and video-paragraph correspondence. To this end, we propose a novel Hierarchical Semantic Correspondence Network (HSCNet), which explores multi-level visual-textual correspondence by learning hierarchical semantic alignment and utilizes dense supervision by grounding diverse levels of queries. Specifically, we develop a hierarchical encoder that encodes the multi-modal inputs into semantics-aligned representations at different levels. To exploit the hierarchical semantic correspondence learned in the encoder for multi-level supervision, we further design a hierarchical decoder that progressively performs finer grounding for lower-level queries conditioned on higher-level semantics. Extensive experiments demonstrate the effectiveness of HSCNet and our method significantly outstrips the state-of-the-arts on two challenging benchmarks, i.e., ActivityNet-Captions and TACoS.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderSentenceVideo Grounding

Similar Papers 제목 키워드 기반

Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding

2024-11-26 · Mengzhao Wang, Huafeng Li, Yafei Zhang, Jinxing Li 외

Video Paragraph Grounding (VPG) aims to precisely locate the most appropriate moments within a video that are relevant to a given textual paragraph query. However, existing methods typically rely on large-scale annotated…

Contrastive LearningRetrieval

Cross-Modal and Hierarchical Modeling of Video and Text

2018-10-16 · ECCV 2018 9 · Bowen Zhang, Hexiang Hu, Fei Sha

Visual data and text data are composed of information at multiple granularities. A video can describe a complex scene that is composed of multiple clips or shots, where each depicts a semantically coherent event or actio…

Action RecognitionRetrievalTemporal Action LocalizationVideo Captioning+1

Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding

2024-03-18 · CVPR 2024 1 · Chaolei Tan, JianHuang Lai, Wei-Shi Zheng, Jian-Fang Hu

Video Paragraph Grounding (VPG) is an emerging task in video-language understanding, which aims at localizing multiple sentences with semantic relations and temporal order from an untrimmed video. However, existing VPG a…

Multiple Instance Learning

SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses

2024-08-03 · Chaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi 외

Video grounding is a fundamental problem in multimodal content understanding, aiming to localize specific natural language queries in an untrimmed video. However, current video grounding datasets merely focus on simple e…

Natural Language QueriesVideo Grounding

Semi-Supervised Video Paragraph Grounding With Contrastive Encoder

2022-01-01 · CVPR 2022 1 · Xun Jiang, Xing Xu, Jingran Zhang, Fumin Shen 외

Video events grounding aims at retrieving the most relevant moments from an untrimmed video in terms of a given natural language query. Most previous works focus on Video Sentence Grounding (VSG), which localizes the…

SentenceVideo Grounding