paper-with-me

Papers

Moment Detection in Long Tutorial Videos

2023-01-01 · ICCV 2023 1 · Ioana Croitoru, Simion-Vlad Bogolin, Samuel Albanie, Yang Liu, Zhaowen Wang, Seunghyun Yoon, Franck Dernoncourt, Hailin Jin, Trung Bui

Tutorial videos play an increasingly important role in professional development and self-directed education. For users to realise the full benefits of this medium, tutorial videos must be efficiently searchable. In this work, we focus on the task of moment detection, in which the goal is to localise the temporal window where a given event occurs within a given tutorial video. Prior work on moment detection has focused primarily on short videos (typically on videos shorter than three minutes). However, many tutorial videos are substantially longer (stretching to hours in duration), presenting significant challenges for existing moment detection approaches. To study this problem, we propose the first dataset of untrimmed, long-form tutorial videos for the task of Moment Detection called the Behance Moment Detection (BMD) dataset. BMD videos have an average duration of over one hour and are characterised by slowly evolving visual content and wide-ranging dialogue. To meet the unique challenges of this dataset, we propose a new framework, LongMoment-DETR, and demonstrate that it outperforms strong baselines. Additionally, we introduce a variation of the dataset that contains YouTube Chapter annotations and show that the features obtained by our framework can be successfully used to boost the performance on the task of chapter detection. Code and data can be found at https://github.com/ioanacroi/longmoment-detr.

📄 PDF Abstract BibTeX

Code (1)

ioanacroi/longmoment-detr 공식 구현

Similar Papers 제목 키워드 기반

Temporal Sentence Grounding in Videos: A Survey and Future Directions

2022-01-20 · Hao Zhang, Aixin Sun, Wei Jing, Joey Tianyi Zhou

Temporal sentence grounding in videos (TSGV), \aka natural language video localization (NLVL) or video moment retrieval (VMR), aims to retrieve a temporal moment that semantically corresponds to a language query from an …

Moment RetrievalRetrievalSentenceTemporal Sentence Grounding

D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX Matching

2024-08-23 · Jingyu Liu, Minquan Wang, Ye Ma, Bo wang 외

Videos showcasing specific products are increasingly important for E-commerce. Key moments naturally exist as the first appearance of a specific product, presentation of its distinctive features, the presence of a buying…

Highlight DetectionMoment Retrieval

Screencast Tutorial Video Understanding

2020-06-01 · CVPR 2020 6 · Kunpeng Li, Chen Fang, Zhaowen Wang, Seokhwan Kim 외

Screencast tutorials are videos created by people to teach how to use software applications or demonstrate procedures for accomplishing tasks. It is very popular for both novice and experienced users to learn new skills,…

object-detectionObject DetectionRetrievalVideo Captioning+2

RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos

2023-12-11 · Tanveer Hannan, Md Mohaiminul Islam, Thomas Seidl, Gedas Bertasius

Locating specific moments within long videos (20-120 minutes) presents a significant challenge, akin to finding a needle in a haystack. Adapting existing short video (5-30 seconds) grounding methods to this problem yield…

Natural Language Moment RetrievalNatural Language QueriesRetrievalText Retrieval+1

A Deep Learning Approach for Multimodal Deception Detection

2018-03-01 · Gangeshwar Krishnamurthy, Navonil Majumder, Soujanya Poria, Erik Cambria

Automatic deception detection is an important task that has gained momentum in computational linguistics due to its potential applications. In this paper, we propose a simple yet tough to beat multi-modal neural model fo…

Deception DetectionDeep Learning