paper-with-me

Papers

Semantic Frame Aggregation-based Transformer for Live Video Comment Generation

2025-10-30 · Anam Fatima, Yi Yu, Janak Kapuriya, Julien Lalanne, Jainendra Shukla arxiv

Live commenting on video streams has surged in popularity on platforms like Twitch, enhancing viewer engagement through dynamic interactions. However, automatically generating contextually appropriate comments remains a challenging and exciting task. Video streams can contain a vast amount of data and extraneous content. Existing approaches tend to overlook an important aspect of prioritizing video frames that are most relevant to ongoing viewer interactions. This prioritization is crucial for producing contextually appropriate comments. To address this gap, we introduce a novel Semantic Frame Aggregation-based Transformer (SFAT) model for live video comment generation. This method not only leverages CLIP's visual-text multimodal knowledge to generate comments but also assigns weights to video frames based on their semantic relevance to ongoing viewer conversation. It employs an efficient weighted sum of frames technique to emphasize informative frames while focusing less on irrelevant ones. Finally, our comment decoder with a cross-attention mechanism that attends to each modality ensures that the generated comment reflects contextual cues from both chats and video. Furthermore, to address the limitations of existing datasets, which predominantly focus on Chinese-language content with limited video categories, we have constructed a large scale, diverse, multimodal English video comments dataset. Extracted from Twitch, this dataset covers 11 video categories, totaling 438 hours and 3.2 million comments. We demonstrate the effectiveness of our SFAT model by comparing it to existing methods for generating comments from live video and ongoing dialogue contexts.

📄 PDF Abstract BibTeX arXiv:2510.26978

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Semantic Segmentation on VSPW Dataset through Aggregation of Transformer Models

2021-09-03 · Zixuan Chen, Junhong Zou, Xiaotao Wang

Semantic segmentation is an important task in computer vision, from which some important usage scenarios are derived, such as autonomous driving, scene parsing, etc. Due to the emphasis on the task of video semantic segm…

Autonomous DrivingScene ParsingSegmentationSemantic Segmentation+1

Temporal Context Aggregation for Video Retrieval with Contrastive Learning

2020-08-04 · Jie Shao, Xin Wen, Bingchen Zhao, xiangyang xue

The current research focus on Content-Based Video Retrieval requires higher-level video representation describing the long-range semantic dependencies of relevant incidents, events, etc. However, existing methods commonl…

Contrastive LearningRepresentation LearningRetrievalVideo Retrieval

COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning

2020-11-01 · NeurIPS 2020 12 · Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, Thomas Brox

Many real-world video-text tasks involve different levels of granularity, such as frames and words, clip and sentences or videos and paragraphs, each with distinct semantics. In this paper, we propose a Cooperative hiera…

Cross-Modal RetrievalRepresentation LearningSentenceVideo Captioning+2

Sentiment-oriented Transformer-based Variational Autoencoder Network for Live Video Commenting

2024-04-19 · Fengyi Fu, Shancheng Fang, Weidong Chen, Zhendong Mao

Automatic live video commenting is with increasing attention due to its significance in narration generation, topic explanation, etc. However, the diverse sentiment consideration of the generated comments is missing from…

Diversity

Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation

2023-05-25 · Shilin Yan, Renrui Zhang, Ziyu Guo, Wenchao Chen 외

Recently, video object segmentation (VOS) referred by multi-modal signals, e.g., language and audio, has evoked increasing attention in both industry and academia. It is challenging for exploring the semantic alignment w…

ObjectReferring Expression SegmentationReferring Video Object SegmentationSemantic Segmentation+2