paper-with-me

홈 › Papers

Video Ads Content Structuring by Combining Scene Confidence Prediction and Tagging

2021-08-20 · Tomoyuki Suzuki, Antonio Tejero-de-Pablos

Video ads segmentation and tagging is a challenging task due to two main reasons: (1) the video scene structure is complex and (2) it includes multiple modalities (e.g., visual, audio, text.). While previous work focuses mostly on activity videos (e.g. "cooking", "sports"), it is not clear how they can be leveraged to tackle the task of video ads content structuring. In this paper, we propose a two-stage method that first provides the boundaries of the scenes, and then combines a confidence score for each segmented scene and the tag classes predicted for that scene. We provide extensive experimental results on the network architectures and modalities used for the proposed method. Our combined method improves the previous baselines on the challenging "Tencent Advertisement Video" dataset.

📄 PDF Abstract BibTeX arXiv:2108.09215

Code (0)

등록된 구현이 없습니다.

Tasks

TAG

Similar Papers 제목 키워드 기반

Multi-modal Representation Learning for Video Advertisement Content Structuring

2021-09-04 · Daya Guo, Zhaoyang Zeng

Video advertisement content structuring aims to segment a given video advertisement and label each segment on various dimensions, such as presentation form, scene, and style. Different from real-life videos, video advert…

Representation LearningRe-RankingVideo Understanding

A Multimodal Framework for Video Ads Understanding

2021-08-29 · Zejia Weng, Lingchen Meng, Rui Wang, Zuxuan Wu 외

There is a growing trend in placing video advertisements on social platforms for online marketing, which demands automatic approaches to understand the contents of advertisements effectively. Taking the 2021 TAAC competi…

MarketingOptical Character RecognitionOptical Character Recognition (OCR)Scene Segmentation+3

ARC-Chapter: Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries

2025-11-18 · Junfu Pu, Teng Wang, Yixiao Ge, Yuying Ge 외 arxiv

The proliferation of hour-long videos (e.g., lectures, podcasts, documentaries) has intensified demand for efficient content structuring. However, existing approaches are constrained by small-scale training with annotati…

Dense Video CaptioningSemantic SimilarityVideo Chaptering

Fuse: In-Situ Sensemaking Support in the Browser

2022-08-31 · Andrew Kuznetsov, Joseph Chee Chang, Nathan Hahn, Napol Rachatasumrit 외

People spend a significant amount of time trying to make sense of the internet, collecting content from a variety of sources and organizing it to make decisions and achieve their goals. While humans are able to fluidly i…

Friction

Automatic Funny Scene Extraction from Long-form Cinematic Videos

2026-02-17 · Sibendu Paul, Haotian Jiang, Caren Chen arxiv

Automatically extracting engaging and high-quality humorous scenes from cinematic titles is pivotal for creating captivating video previews and snackable content, boosting user engagement on streaming platforms. Long-for…

Scene Segmentation