paper-with-me

홈 › Papers

MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video Segmentation

2023-08-22 · ICCV 2023 1 · Najmeh Sadoughi, Xinyu Li, Avijit Vajpayee, David Fan, Bing Shuai, Hector Santos-Villalobos, Vimal Bhat, Rohith MV

Previous research has studied the task of segmenting cinematic videos into scenes and into narrative acts. However, these studies have overlooked the essential task of multimodal alignment and fusion for effectively and efficiently processing long-form videos (>60min). In this paper, we introduce Multimodal alignmEnt aGgregation and distillAtion (MEGA) for cinematic long-video segmentation. MEGA tackles the challenge by leveraging multiple media modalities. The method coarsely aligns inputs of variable lengths and different modalities with alignment positional encoding. To maintain temporal synchronization while reducing computation, we further introduce an enhanced bottleneck fusion layer which uses temporal alignment. Additionally, MEGA employs a novel contrastive loss to synchronize and transfer labels across modalities, enabling act segmentation from labeled synopsis sentences on video shots. Our experimental results show that MEGA outperforms state-of-the-art methods on MovieNet dataset for scene segmentation (with an Average Precision improvement of +1.19%) and on TRIPOD dataset for act segmentation (with a Total Agreement improvement of +5.51%)

📄 PDF Abstract BibTeX arXiv:2308.11185

Code (0)

등록된 구현이 없습니다.

Tasks

Scene SegmentationSegmentationVideo SegmentationVideo Semantic Segmentation

Similar Papers 제목 키워드 기반

Customized Visual Storytelling with Unified Multimodal LLMs

2026-03-29 · Wei-Hua Li, Cheng Sun, Chu-Song Chen arxiv

Multimodal story customization aims to generate coherent story flows conditioned on textual descriptions, reference identity images, and shot types. While recent progress in story generation has shown promising results, …

Visual StorytellingStory Generation

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation

2025-07-03 · Feizhen Huang, Yu Wu, Yutian Lin, Bo Du arxiv

Video-to-Audio (V2A) Generation achieves significant progress and plays a crucial role in film and video post-production. However, current methods overlook the cinematic language, a critical component of artistic express…

Audio Generation

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation

2026-01-25 · Chenyu Mu, Xin He, Qu Yang, Wanshun Chen 외 arxiv

Recent advances in video generation have produced models capable of synthesizing stunning visual content from simple text prompts. However, these models struggle to generate long-form, coherent narratives from high-level…

Video Generation

Automatic Funny Scene Extraction from Long-form Cinematic Videos

2026-02-17 · Sibendu Paul, Haotian Jiang, Caren Chen arxiv

Automatically extracting engaging and high-quality humorous scenes from cinematic titles is pivotal for creating captivating video previews and snackable content, boosting user engagement on streaming platforms. Long-for…

Scene Segmentation

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

2026-07-27 · Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi 외 hf

Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained…

Video Generation