paper-with-me

Papers

Sentence Guided Temporal Modulation for Dynamic Video Thumbnail Generation

2020-08-31 · Mrigank Rochan, Mahesh Kumar Krishna Reddy, Yang Wang

We consider the problem of sentence specified dynamic video thumbnail generation. Given an input video and a user query sentence, the goal is to generate a video thumbnail that not only provides the preview of the video content, but also semantically corresponds to the sentence. In this paper, we propose a sentence guided temporal modulation (SGTM) mechanism that utilizes the sentence embedding to modulate the normalized temporal activations of the video thumbnail generation network. Unlike the existing state-of-the-art method that uses recurrent architectures, we propose a non-recurrent framework that is simple and allows much more parallelization. Extensive experiments and analysis on a large-scale dataset demonstrate the effectiveness of our framework.

📄 PDF Abstract BibTeX arXiv:2008.13362

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceSentence EmbeddingSentence-Embedding

Similar Papers 제목 키워드 기반

Semantic Conditioned Dynamic Modulation for Temporal Sentence Grounding in Videos

2019-10-31 · NeurIPS 2019 12 · Yitian Yuan, Lin Ma, Jingwen Wang, Wei Liu 외

Temporal sentence grounding in videos aims to detect and localize one target video segment, which semantically corresponds to a given sentence. Existing methods mainly tackle this task via matching and aligning semantics…

SentenceTemporal Sentence Grounding

Reasoning Step-by-Step: Temporal Sentence Localization in Videos via Deep Rectification-Modulation Network

2020-12-01 · COLING 2020 8 · Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou

Temporal sentence localization in videos aims to ground the best matched segment in an untrimmed video according to a given sentence query. Previous works in this field mainly rely on attentional frameworks to align the …

Sentence

Language Guided Networks for Cross-modal Moment Retrieval

2020-06-18 · Kun Liu, Huadong Ma, Chuang Gan

We address the challenging task of cross-modal moment retrieval, which aims to localize a temporal segment from an untrimmed video described by a natural language query. It poses great challenges over the proper semantic…

Moment RetrievalRetrievalSentenceSentence Embedding+1

IDAG-Edit: Multi-Object Video Editing via Instance-Decoupled Attention and Guidance

2026-06-20 · Yuan-Zhih Lin, Huu-Thang Nguyen, Huu-Phu Do, Hong-Han Shuai 외 arxiv

Diffusion-based video editing has made significant progress; however, achieving precise and temporally consistent object-level control, especially in multi-object scenarios, remains challenging due to attention leakage, …

Proposal-free Temporal Moment Localization of a Natural-Language Query in Video using Guided Attention

2019-08-20 · Cristian Rodriguez-Opazo, Edison Marrese-Taylor, Fatemeh Sadat Saleh, Hongdong Li 외

This paper studies the problem of temporal moment localization in a long untrimmed video using natural language as the query. Given an untrimmed video and a sentence as the query, the goal is to determine the starting, a…

Sentence