paper-with-me

홈 › Papers

Text-based Localization of Moments in a Video Corpus

2020-08-20 · Sudipta Paul, Niluthpol Chowdhury Mithun, Amit K. Roy-Chowdhury

Prior works on text-based video moment localization focus on temporally grounding the textual query in an untrimmed video. These works assume that the relevant video is already known and attempt to localize the moment on that relevant video only. Different from such works, we relax this assumption and address the task of localizing moments in a corpus of videos for a given sentence query. This task poses a unique challenge as the system is required to perform: (i) retrieval of the relevant video where only a segment of the video corresponds with the queried sentence, and (ii) temporal localization of moment in the relevant video based on sentence query. Towards overcoming this challenge, we propose Hierarchical Moment Alignment Network (HMAN) which learns an effective joint embedding space for moments and sentences. In addition to learning subtle differences between intra-video moments, HMAN focuses on distinguishing inter-video global semantic concepts based on sentence queries. Qualitative and quantitative results on three benchmark text-based video moment retrieval datasets - Charades-STA, DiDeMo, and ActivityNet Captions - demonstrate that our method achieves promising performance on the proposed task of temporal localization of moments in a corpus of videos.

📄 PDF Abstract BibTeX arXiv:2008.08716

Code (0)

등록된 구현이 없습니다.

Tasks

Moment RetrievalRetrievalSentenceTemporal Localization

Similar Papers 제목 키워드 기반

Finding Moments in Video Collections Using Natural Language

2019-07-30 · Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem 외

We introduce the task of retrieving relevant video moments from a large corpus of untrimmed, unsegmented videos given a natural language query. Our task poses unique challenges as a system must efficiently identify both …

Moment RetrievalRe-RankingRetrievalTemporal Localization+1

Multi-video Moment Ranking with Multimodal Clue

2023-01-29 · Danyang Hou, Liang Pang, Yanyan Lan, HuaWei Shen 외

Video corpus moment retrieval~(VCMR) is the task of retrieving a relevant video moment from a large corpus of untrimmed videos via a natural language query. State-of-the-art work for VCMR is based on two-stage method. In…

Moment RetrievalRetrievalVideo Corpus Moment Retrieval

Learning to Localize Actions from Moments

2020-08-31 · ECCV 2020 8 · Fuchen Long, Ting Yao, Zhaofan Qiu, Xinmei Tian 외

With the knowledge of action moments (i.e., trimmed video clips that each contains an action instance), humans could routinely localize an action temporally in an untrimmed video. Nevertheless, most practical methods sti…

Action LocalizationTransfer Learning

Multi-Scale 2D Temporal Adjacent Networks for Moment Localization with Natural Language

2020-12-04 · Songyang Zhang, Houwen Peng, Jianlong Fu, Yijuan Lu 외

We address the problem of retrieving a specific moment from an untrimmed video by natural language. It is a challenging problem because a target moment may take place in the context of other temporal moments in the untri…

Progressive Localization Networks for Language-based Moment Localization

2021-02-02 · Qi Zheng, Jianfeng Dong, Xiaoye Qu, Xun Yang 외

This paper targets the task of language-based video moment localization. The language-based setting of this task allows for an open set of target activities, resulting in a large variation of the temporal lengths of vide…