paper-with-me

홈 › Papers

Multi-Modal Video Topic Segmentation with Dual-Contrastive Domain Adaptation

2023-11-30 · Linzi Xing, Quan Tran, Fabian Caba, Franck Dernoncourt, Seunghyun Yoon, Zhaowen Wang, Trung Bui, Giuseppe Carenini

Video topic segmentation unveils the coarse-grained semantic structure underlying videos and is essential for other video understanding tasks. Given the recent surge in multi-modal, relying solely on a single modality is arguably insufficient. On the other hand, prior solutions for similar tasks like video scene/shot segmentation cater to short videos with clear visual shifts but falter for long videos with subtle changes, such as livestreams. In this paper, we introduce a multi-modal video topic segmenter that utilizes both video transcripts and frames, bolstered by a cross-modal attention mechanism. Furthermore, we propose a dual-contrastive learning framework adhering to the unsupervised domain adaptation paradigm, enhancing our model's adaptability to longer, more semantically complex videos. Experiments on short and long video corpora demonstrate that our proposed solution, significantly surpasses baseline methods in terms of both accuracy and transferability, in both intra- and cross-domain settings.

📄 PDF Abstract BibTeX arXiv:2312.00220

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningDomain AdaptationSegmentationUnsupervised Domain AdaptationVideo Understanding

Similar Papers 제목 키워드 기반

Multimodal Fusion and Coherence Modeling for Video Topic Segmentation

2024-08-01 · Hai Yu, Chong Deng, Qinglin Zhang, Jiaqing Liu 외

The video topic segmentation (VTS) task segments videos into intelligible, non-overlapping topics, facilitating efficient comprehension of video content and quick access to specific content. VTS is also critical to vario…

Contrastive LearningMixture-of-ExpertsScene SegmentationSegmentation+1

VSTAR: A Video-grounded Dialogue Dataset for Situated Semantic Understanding with Scene and Topic Transitions

2023-05-30 · Yuxuan Wang, Zilong Zheng, Xueliang Zhao, Jinpeng Li 외

Video-grounded dialogue understanding is a challenging problem that requires machine to perceive, parse and reason over situated semantics extracted from weakly aligned video and dialogues. Most existing benchmarks treat…

Dialogue GenerationDialogue UnderstandingScene SegmentationSegmentation

Reading Between the Waves: Robust Topic Segmentation Using Inter-Sentence Audio Features

2026-02-06 · Steffen Freisinger, Philipp Seeberger, Tobias Bocklet, Korbinian Riedhammer arxiv

Spoken content, such as online videos and podcasts, often spans multiple topics, which makes automatic topic segmentation essential for user navigation and downstream applications. However, current methods do not fully l…

MMTM: Tri-Modal Topic Modeling for Long-Form Video via Similarity-Gated Fusion

2026-05-28 · Ali Abusaleh, Bhuvanesh Verma, Alexander Mehler arxiv

We introduce MMTM, a modular pipeline for topic discovery in long-form video that integrates speech recognition, audio and visual embeddings, and BERTopic clustering through a deterministic similarity-gated fusion. Evalu…

Speech Recognition

LiveSeg: Unsupervised Multimodal Temporal Segmentation of Long Livestream Videos

2022-10-12 · JieLin Qiu, Franck Dernoncourt, Trung Bui, Zhaowen Wang 외

Livestream videos have become a significant part of online learning, where design, digital marketing, creative painting, and other skills are taught by experienced experts in the sessions, making them valuable materials.…

MarketingSegmentation