paper-with-me

홈 › Papers

Boundary Proposal Network for Two-Stage Natural Language Video Localization

2021-03-15 · Shaoning Xiao, Long Chen, Songyang Zhang, Wei Ji, Jian Shao, Lu Ye, Jun Xiao

We aim to address the problem of Natural Language Video Localization (NLVL)-localizing the video segment corresponding to a natural language description in a long and untrimmed video. State-of-the-art NLVL methods are almost in one-stage fashion, which can be typically grouped into two categories: 1) anchor-based approach: it first pre-defines a series of video segment candidates (e.g., by sliding window), and then does classification for each candidate; 2) anchor-free approach: it directly predicts the probabilities for each video frame as a boundary or intermediate frame inside the positive segment. However, both kinds of one-stage approaches have inherent drawbacks: the anchor-based approach is susceptible to the heuristic rules, further limiting the capability of handling videos with variant length. While the anchor-free approach fails to exploit the segment-level interaction thus achieving inferior results. In this paper, we propose a novel Boundary Proposal Network (BPNet), a universal two-stage framework that gets rid of the issues mentioned above. Specifically, in the first stage, BPNet utilizes an anchor-free model to generate a group of high-quality candidate video segments with their boundaries. In the second stage, a visual-language fusion layer is proposed to jointly model the multi-modal interaction between the candidate and the language query, followed by a matching score rating layer that outputs the alignment score for each candidate. We evaluate our BPNet on three challenging NLVL benchmarks (i.e., Charades-STA, TACoS and ActivityNet-Captions). Extensive experiments and ablative studies on these datasets demonstrate that the BPNet outperforms the state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2103.08109

Code (0)

등록된 구현이 없습니다.

Tasks

Vocal Bursts Valence Prediction

Methods 이 논문이 사용한 방법론

Error Back Proragation Network 설명 없음

Similar Papers 제목 키워드 기반

Natural Language Video Localization with Learnable Moment Proposals

2021-09-22 · EMNLP 2021 11 · Shaoning Xiao, Long Chen, Jian Shao, Yueting Zhuang 외

Given an untrimmed video and a natural language query, Natural Language Video Localization (NLVL) aims to identify the video moment described by the query. To address this task, existing methods can be roughly grouped in…

CMSN: Continuous Multi-stage Network and Variable Margin Cosine Loss for Temporal Action Proposal Generation

2019-11-14 · Yushuai Hu, Yaochu Jin, Runhua Li, Xiangxiang Zhang

Accurately locating the start and end time of an action in untrimmed videos is a challenging task. One of the important reasons is the boundary of action is not highly distinguishable, and the features around the boundar…

Temporal Action Proposal Generation

Structured Multi-Level Interaction Network for Video Moment Localization via Language Query

2021-06-19 · CVPR 2021 1 · Hao Wang, Zheng-Jun Zha, Liang Li, Dong Liu 외

We address the problem of localizing a specific moment described by a natural language query. Existing works interact the query with either video frame or moment proposal, and neglect the inherent structure of moment…

Sentence

Locate and Label: A Two-stage Identifier for Nested Named Entity Recognition

2021-05-14 · ACL 2021 5 · Yongliang Shen, Xinyin Ma, Zeqi Tan, Shuai Zhang 외

Named entity recognition (NER) is a well-studied task in natural language processing. Traditional NER research only deals with flat entities and ignores nested entities. The span-based methods treat entity recognition as…

Chinese Named Entity Recognitionnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+2

Look Closer to Ground Better: Weakly-Supervised Temporal Grounding of Sentence in Video

2020-01-25 · Zhenfang Chen, Lin Ma, Wenhan Luo, Peng Tang 외

In this paper, we study the problem of weakly-supervised temporal grounding of sentence in video. Specifically, given an untrimmed video and a query sentence, our goal is to localize a temporal segment in the video that …

Sentence