paper-with-me

홈 › Papers

TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization

2026-03-01 · Sumin Kim, Hyemin Jeong, Mingu Kang, Yejin Kim, Yoori Oh, Joonseok Lee arxiv

The exponential growth of video content necessitates effective video summarization to efficiently extract key information from long videos. However, current approaches struggle to fully comprehend complex videos, primarily because they employ static or modality-agnostic fusion strategies. These methods fail to account for the dynamic, frame-dependent variations in modality saliency inherent in video data. To overcome these limitations, we propose TripleSumm, a novel architecture that adaptively weights and fuses the contributions of visual, text, and audio modalities at the frame level. Furthermore, a significant bottleneck for research into multimodal video summarization has been the lack of comprehensive benchmarks. Addressing this bottleneck, we introduce MoSu (Most Replayed Multimodal Video Summarization), the first large-scale benchmark that provides all three modalities. Extensive experiments demonstrate that TripleSumm achieves state-of-the-art performance, outperforming existing methods by a significant margin on four benchmarks, including MoSu. Our code and dataset are available at https://github.com/smkim37/TripleSumm.

📄 PDF Abstract BibTeX arXiv:2603.01169

Code (0)

등록된 구현이 없습니다.

Tasks

Video Summarization

Similar Papers 제목 키워드 기반

Unleashing the Power of Imbalanced Modality Information for Multi-modal Knowledge Graph Completion

2024-02-22 · Yichi Zhang, Zhuo Chen, Lei Liang, Huajun Chen 외

Multi-modal knowledge graph completion (MMKGC) aims to predict the missing triples in the multi-modal knowledge graphs by incorporating structural, visual, and textual information of entities into the discriminant models…

Knowledge Graph CompletionKnowledge GraphsMulti-modal Knowledge Graph

M2FNet: Multi-modal Fusion Network for Emotion Recognition in Conversation

2022-06-05 · Vishal Chudasama, Purbayan Kar, Ashish Gudmalwar, Nirmesh Shah 외

Emotion Recognition in Conversations (ERC) is crucial in developing sympathetic human-machine interaction. In conversational videos, emotion can be present in multiple modalities, i.e., audio, video, and transcript. Howe…

Emotion RecognitionEmotion Recognition in ConversationTriplet

Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos

2026-05-25 · Bonan Ding, Umair Nawaz, Ufaq Khan, Abdelrahman M. Shaker 외 arxiv

Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evidence. In such a scenario, uniform fusio…

Visual Reasoning

AdaFuse: Adaptive Multimodal Fusion for Lung Cancer Risk Prediction via Reinforcement Learning

2026-01-30 · Chongyu Qu, Zhengyi Lu, Yuxiang Lai, Thomas Z. Li 외 arxiv

Multimodal fusion has emerged as a promising paradigm for disease diagnosis and prognosis, integrating complementary information from heterogeneous data sources such as medical images, clinical records, and radiology rep…

Reinforcement Learning

OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding

2025-04-15 · Dianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu 외

In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff, aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats al…

Semantic SegmentationVideo GenerationVideo Understanding