paper-with-me

홈 › Papers

Automatic Detection of Intro and Credits in Video using CLIP and Multihead Attention

2025-04-13 · Vasilii Korolkov, Andrey Yanchenko

Detecting transitions between intro/credits and main content in videos is a crucial task for content segmentation, indexing, and recommendation systems. Manual annotation of such transitions is labor-intensive and error-prone, while heuristic-based methods often fail to generalize across diverse video styles. In this work, we introduce a deep learning-based approach that formulates the problem as a sequence-to-sequence classification task, where each second of a video is labeled as either "intro" or "film." Our method extracts frames at a fixed rate of 1 FPS, encodes them using CLIP (Contrastive Language-Image Pretraining), and processes the resulting feature representations with a multihead attention model incorporating learned positional encoding. The system achieves an F1-score of 91.0%, Precision of 89.0%, and Recall of 97.0% on the test set, and is optimized for real-time inference, achieving 11.5 FPS on CPU and 107 FPS on high-end GPUs. This approach has practical applications in automated content indexing, highlight detection, and video summarization. Future work will explore multimodal learning, incorporating audio features and subtitles to further enhance detection accuracy.

📄 PDF Abstract BibTeX arXiv:2504.09738

Code (0)

등록된 구현이 없습니다.

Tasks

CPUHighlight DetectionRecommendation SystemsVideo Summarization

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding

2025-07-03 · Xiangfeng Wang, Xiao Li, Yadong Wei, Xueyu Song 외 arxiv

The rapid growth of online video content, especially on short video platforms, has created a growing demand for efficient video editing techniques that can condense long-form videos into concise and engaging clips. Exist…

Highlight Detection

RodEpil: A Video Dataset of Laboratory Rodents for Seizure Detection and Benchmark Evaluation

2025-11-13 · Daniele Perlo, Vladimir Despotovic, Selma Boudissa, Sang-Yoon Kim 외 arxiv

We introduce a curated video dataset of laboratory rodents for automatic detection of convulsive events. The dataset contains short (10~s) top-down and side-view video clips of individual rodents, labeled at clip level a…

Seizure Detection

Unsupervised Multi-stream Highlight detection for the Game "Honor of Kings"

2019-10-14 · Li Wang, Zixun Sun, Wentao Yao, Hui Zhan 외

With the increasing popularity of E-sport live, Highlight Flashback has been a critical functionality of live platforms, which aggregates the overall exciting fighting scenes in a few seconds. In this paper, we introduce…

Highlight Detection

CLIP-VAD: Exploiting Vision-Language Models for Voice Activity Detection

2024-10-18 · Andrea Appiani, Cigdem Beyan

Voice Activity Detection (VAD) is the process of automatically determining whether a person is speaking and identifying the timing of their speech in an audiovisual data. Traditionally, this task has been tackled by proc…

Action DetectionActivity DetectionPrompt Engineering

Towards Unsupervised Familiar Scene Recognition in Egocentric Videos

2019-05-10 · Estefania Talavera, Nicolai Petkov, Petia Radeva

Nowadays, there is an upsurge of interest in using lifelogging devices. Such devices generate huge amounts of image data; consequently, the need for automatic methods for analyzing and summarizing these data is drastical…

Scene Recognition